MWFAI Middleware Factory / Artificial Intelligence

Agent Qualification

You have agents. You don't have numbers.

Which model, which tools, which policy — and what each one costs you per task. A reserved envelope of days over two to four weeks, with a budget ceiling agreed up front and no technology to adopt first.

Teams pick a model because it benchmarked well on someone else's tasks, expose forty tools because the server offers forty, and find out months later that half the failures come from the tool surface rather than the model.

We run your tasks against controlled variants — model, instructions, exposed tools, agent policy — and report accuracy, failures, round trips, tokens and duration for each. Same task, same scoring, one variable at a time.

What measuring finds, in practice

Three findings from our own qualification runs, to show the kind of thing that only appears when you measure:

  • a command-line runner that silently substituted a different model's profile from the one it was told to use;
  • a call ceiling that bounded the main agent but not the sub-agent it spawned;
  • a reduced tool surface that kept the same accuracy for 3.3× fewer tokens.

None of these are visible from the outside. All three change what you should ship.

The form, stated honestly

Our harness has hundreds of recorded runs — the proof dossier gives the exact count and how it is counted — with versioned policies and deterministic scoring. Today it measures agents operating a Generic System model, which makes the comparative form immediately available: your stack against a governed model, on your tasks.

Qualifying your own stack end to end is scoped work — the task format and the scoring have to travel outside our model. We would rather tell you that than discover it at week three.

Pick three tasks that matter.

That is enough to produce a report you can act on.