Use Cases
Teams bring Jetty a job written in plain language, pick an agent and model, and let it run at scale: hundreds of trajectories in parallel, every configuration saved, every result scored by evals. What comes back is organized for investigation, so the next run is better than the last.
Each story covers the situation, what was set up on Jetty, and what came out of it.
Nous ResearchML Engineer / Researcher
A small eval beats a big rewrite: improving skills in development
How Morgane Moss at Nous Research used Jetty to run hundreds of long agent trajectories in parallel and optimize a skill through cheap, repeatable evals.
Morgane Moss · ML Engineering, Nous Research
Read the storyThe AI workload metrics
- Runs
- +15k
- Avg. evaluations per run
- 8
- Run time configs
- 30–90 min
- On Jetty
- Author
- Run
- Investigate
- Improve