NeoCognition’s ApprenticeBench tests whether an AI agent can learn a new job and keep working over a long sequence of tasks, rather than completing only one benchmark task.
How it works
The agent joins a virtual construction company in an accounts-payable role. It receives a six-month invoice archive, an internal handbook, and an ERP guide. During the first month it receives detailed feedback; later, feedback becomes infrequent. The benchmark contains 100 tasks reviewed by professional construction-industry accountants.
Results
According to NeoCognition, Fable 5.1 completed 72 percent of the tasks, GPT-6 Astra completed 68 percent, and the best of two human testers reached 51 percent. Most other models remained around 20–25 percent or lower.
Weaker models rarely returned to the invoice archive or their own notes and degraded toward the end of the sequence.
What learning adds
Without history or feedback, Fable 5 solved 11 out of 100 tasks. Adding the archive raised the result to 31, manager feedback raised it to 35, and using both sources produced 43.
Memory is not the same as learning. An agent must know when to retrieve information, turn feedback into a working rule, and preserve improvements across later tasks.
Why it matters
ApprenticeBench is closer to a real deployment question than a one-shot benchmark: can an agent absorb context, use history, avoid repeating mistakes, and remain stable over time?

