Research
Jetty works with research labs and standards groups on how AI systems get evaluated: how benchmarks are packaged, how evaluations are made reproducible, and how models are held accountable once they leave the lab.
These are the papers and whitepapers we have contributed to, newest first.
· arXiv preprint
Croissant Tasks: A Metadata Format for Reproducible Machine Learning Evaluations
A declarative metadata format that separates an evaluation task from its solution, so agents can rebuild a benchmark’s pipeline from the spec instead of replicating its source code.
- Authors
- Omar Benjelloun, Leonardo Martins Bianco, Isabelle Guyon, Thanh Gia Hieu Khuong, Jonathan Lebensold, Sebastian Lobentanzer, Luis Oala, Benedictus Kent Rachmat, Ihsan Ullah, Peyman Vahidi, Joaquin Vanschoren
- Affiliations
- Google DeepMind, ChaLearn, Université Paris-Saclay, Jetty, Mila – Quebec AI Institute, Helmholtz Munich, German Center for Diabetes Research, Technical University of Munich, Helmholtz AI, Brickroad, Eindhoven University of Technology
· arXiv preprint
CUBE: A Standard for Unifying Agent Benchmarks
A protocol standard built on MCP and Gym so an agent benchmark can be wrapped once and used by any compliant platform for evaluation, RL training or data generation.
- Authors
- Alexandre Lacoste, Nicolas Gontier, Oleh Shliazhko, Aman Jaiswal, Kusha Sareen, Shailesh Nanisetty, Joan Cabezas, Manuel Del Verme, Omar G. Younis, Simone Baratta, Matteo Avalle, Imene Kerboua, Xing Han Lù, Elron Bandel, Michal Shmueli-Scheuer, Asaf Yehudai, Leshem Choshen, Jonathan Lebensold, Sean Hughes, Massimo Caccia, Alexandre Drouin, Siva Reddy, Tao Yu, Yu Su, Graham Neubig, Dawn Song
- Affiliations
- ServiceNow AI Research, Silverstream.ai, Dalhousie University, IBM Research, Jetty, McGill University, Mila, The University of Hong Kong, The Ohio State University, Carnegie Mellon University, UC Berkeley
· Whitepaper
AI Beyond Metrics: Insights from SMASH 2025
Takeaways from the inaugural Symposium on Model Accountability, Sustainability and Healthcare, hosted at Mila, on judging AI by trust, governance and sustainability rather than benchmark scores alone.
- Authors
- Göktuğ BenderForeword by Jonathan Lebensold (Jetty), SMASH co-organizer.
- Partners
- Akinox, CDSI (McGill University), CHAI, CIFAR, Google Sustainability, The Google School for Leaders, Jetty, Mila, MLCommons, OBVIA