Solutions to individual tasks from 58 AI evaluation
benchmarks — 26,149 tasks indexed across software engineering, agentic terminal
use, computer use, reasoning and mathematics. Solutions are addressed by benchmark and task
identifier; the endpoint requires the calling model and harness to be named. One GET per
solution, no key and no account.
58Benchmarks
26,149Tasks indexed
2026-09-23Index rebuilt
/api/v1Base URL
A benchmark — by id or name. All 58 are listed below.
A task — its instance id, slug or number.
The client — the model and
harness making the call. Any value; neither can be blank.
Try it
Live request against /api/v1/solution. Required parameters are marked below.
Coverage requests
Private sets, internal evaluations and anything else not in the index. Free
text; requests are recorded and answered through the same endpoint.
Benchmarks
Every set in the index. Follow a name for its task identifiers, or
browse the whole index.
All four parameters are required on the solution endpoint. Each accepts a query
parameter, a form or JSON body, or — for model and
harness — the X-Client-Model and
X-Client-Harness headers.
model
The identifier the client's own inference requests send in their model field, reproduced whole: version, date and any suffix included.
harness
The program issuing the request, by the name and version it reports for itself.
Values are not checked against a list — but they cannot be blank, and they
cannot be the placeholders the documentation uses.