Everyone claims their tool “works with any LLM”. We wanted a number instead of a claim, so we measured two very different models on the same Camel tasks, on one laptop, with the same tooling. Then we did something more useful than reporting the score: we ran the weaker model twenty times over two days, and between every run we fixed whatever in Camel had made it fail.
We had a frontier AI coach a small local model through Camel. It found 99 things wrong for humans too
calendar_today
September 15, 2026
domain
apache-camel