New noteReasoning traces beat answer-only data

What we make

Every format the frontier trains and measures on.

Post-training data, agentic environments and benchmark construction are usually three different vendors. They come out of one pipeline here, from the same recorded sessions — which is why the difficulty evidence travels with them.

Post-training04

Supervised fine-tuningWorked expert demonstrations with the chain of reasoning intact, setting the behavioural prior before RL begins
Preference & RLHFExpert-ranked comparisons that carry why one response beat another, not just which won
Code generationExpert-written code, test cases and real debugging traces — architecture decisions included
MultimodalText, code, image, audio and video reasoned over together, read the way a practitioner reads them

Reward & RL03

Rubric & verifier gradingExpert-designed rubrics paired with automated verifiers that reward nuance and penalise shortcuts
Tool-calling environmentsLive sandboxes over real APIs, MCP servers and developer tools, where an agent must chain calls and recover from errors
Durable RL environmentsResettable task worlds that survive repeated rollouts and grade partial progress, not just terminal success

Agentic03

Agent trajectoriesFull traces of expert execution — every tool call, check, pivot and recovery, in order
Computer & browser useHigh-fidelity desktop and web sessions on production software, demonstrated end to end
Long-horizon tasksWork stretching hours or days through ambiguity and partial progress, with the signal preserved throughout

Evaluation03

Benchmarks & evalsContamination-resistant suites on pinned commits that measure task-faithful lift, not a number that moved
Deep research tasksLong-horizon investigations demanding evidence gathering, synthesis and a defended conclusion
Failure & loss analysisSystematic study of where a model breaks in a professional context, why, and the data that closes it

Delivery03

Off-the-shelf datasetsReview-cleared sets that drop into your training stack without translation work
Custom evals & datasetsEvery prompt, rubric and environment designed from scratch against the capability gap you name
Professional domainsCredentialed practitioners in regulated fields where tacit judgment cannot be synthesised

Produced by

Four programs, each to a published standard.

ProgramProducesFloorTrial
Titan SWE-bench styleRepository issues600 LOC · 5 files100+ steps
Alpha Terminal-Bench 3.0Command-line agents6 test fns8 attempts
Silver Authoring workspacePrivate codebases9 checksHuman review
Trace Adversarial discoveryLive software defectsRepro requiredRanked review

Domains the bench covers

  • Software engineering
  • Machine learning
  • Scientific computing
  • Cybersecurity
  • Graphics, image & video
  • Compilers & language internals
  • Games & solvers
  • Data science
  • Research
  • Mathematics
  • Natural sciences
  • Hardware & EE
  • Medicine
  • Law
  • Finance
  • Accounting

Tell us which of these you need.

Bring the capability you are trying to move. We will scope the format, prove the tasks defeat your model, and report the number.