Documentation menu

Research

Jetty works with research labs and standards groups on how AI systems get evaluated: how benchmarks are packaged, how evaluations are made reproducible, and how models are held accountable once they leave the lab.

These are the papers and whitepapers we have contributed to, newest first.

  1. · arXiv preprint

    Croissant Tasks: A Metadata Format for Reproducible Machine Learning Evaluations

    A declarative metadata format that separates an evaluation task from its solution, so agents can rebuild a benchmark’s pipeline from the spec instead of replicating its source code.

    Authors
    Omar Benjelloun, Leonardo Martins Bianco, Isabelle Guyon, Thanh Gia Hieu Khuong, Jonathan Lebensold, Sebastian Lobentanzer, Luis Oala, Benedictus Kent Rachmat, Ihsan Ullah, Peyman Vahidi, Joaquin Vanschoren
    Affiliations
    Google DeepMind, ChaLearn, Université Paris-Saclay, Jetty, Mila – Quebec AI Institute, Helmholtz Munich, German Center for Diabetes Research, Technical University of Munich, Helmholtz AI, Brickroad, Eindhoven University of Technology
  2. · arXiv preprint

    CUBE: A Standard for Unifying Agent Benchmarks

    A protocol standard built on MCP and Gym so an agent benchmark can be wrapped once and used by any compliant platform for evaluation, RL training or data generation.

    Authors
    Alexandre Lacoste, Nicolas Gontier, Oleh Shliazhko, Aman Jaiswal, Kusha Sareen, Shailesh Nanisetty, Joan Cabezas, Manuel Del Verme, Omar G. Younis, Simone Baratta, Matteo Avalle, Imene Kerboua, Xing Han Lù, Elron Bandel, Michal Shmueli-Scheuer, Asaf Yehudai, Leshem Choshen, Jonathan Lebensold, Sean Hughes, Massimo Caccia, Alexandre Drouin, Siva Reddy, Tao Yu, Yu Su, Graham Neubig, Dawn Song
    Affiliations
    ServiceNow AI Research, Silverstream.ai, Dalhousie University, IBM Research, Jetty, McGill University, Mila, The University of Hong Kong, The Ohio State University, Carnegie Mellon University, UC Berkeley
  3. · Whitepaper

    AI Beyond Metrics: Insights from SMASH 2025

    Takeaways from the inaugural Symposium on Model Accountability, Sustainability and Healthcare, hosted at Mila, on judging AI by trust, governance and sustainability rather than benchmark scores alone.

    Authors
    Göktuğ BenderForeword by Jonathan Lebensold (Jetty), SMASH co-organizer.
    Partners
    Akinox, CDSI (McGill University), CHAI, CIFAR, Google Sustainability, The Google School for Leaders, Jetty, Mila, MLCommons, OBVIA