
Litmus
VerifiedEvals for humans
Elevator Pitch
Evals for humans
About
Litmus is building the most accurate framework for evaluating and benchmarking human capability, starting with software. AI will compound small differences in human capability into increasingly large differences in what people can accomplish, while making existing static benchmarks obsolete. Software is already there: AI can hill-climb any output-based evaluation, while the ability to direct it is becoming the defining advantage. Litmus applies the same approach we already use for model evals to humans – creating world-like environments, and inspecting trajectory instead of just output. Every knowledge industry will soon face the same problem. We build Litmus to tell you what humans are capable of. We're already helping build frontier technical teams at Mercor, Composio, Neo Scholars, and more.