Notice: Undefined offset: 1 in /var/www/tmr/wp-content/plugins/accelerated-mobile-pages/includes/vendor/amp/includes/utils/class-amp-image-dimension-extractor.php on line 244
Snorkel AI Highlights First Wave of Open Benchmarks Grants Projects - The Malaysian Reserve
Categories: PR Newswire

Snorkel AI Highlights First Wave of Open Benchmarks Grants Projects

SAN FRANCISCO, July 24, 2026 /PRNewswire/ — Snorkel AI today highlighted the first group of projects supported through Open Benchmarks Grants, a $3 million commitment to support open-source datasets, benchmarks, and evaluation research.

Launched in February 2026, Open Benchmarks Grants has received hundreds of applications from researchers, labs, and engineers working to address a growing challenge: AI systems are advancing faster than the field’s ability to rigorously measure their performance on realistic, consequential work.

“From complex environments and huge autonomy horizons to rich, sophisticated outputs, these projects tackle some of the field’s hardest evaluation challenges,” said Fred Sala, a member of the Open Benchmarks Grants steering committee and assistant professor at the University of Wisconsin–Madison. “I’m excited to see the broader research community use, validate, and build on them.”

Open Benchmarks Grants provides selected teams with funding, expert data development support, research and engineering collaboration, and platform resources. Supported projects include:

  • Frontier-Bench (formerly Terminal-Bench 3.0), developed with Laude Institute and the Harbor community, is a harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.
  • Agents’ Last Exam, developed with UC Berkeley RDI and the RDI Foundation, evaluates agents on long-horizon, economically valuable professional workflows. It spans 55 sub-industries and includes more than 1,500 tasks toward a 5,000-task target, sourced and validated by more than 300 industry experts.
  • OSWorld 2.0, developed with XLANG Lab, evaluates computer-use agents on 108 long-horizon workflows across 31 self-hosted web environments and professional desktop applications.
  • Continual Learning Bench, developed with UC Berkeley SkyLab and the University of Wisconsin–Madison, measures whether agents genuinely improve across sequential, stateful tasks.
  • SlopCode Bench, developed with the University of Wisconsin–Madison, measures how code quality degrades as coding agents repeatedly modify and extend their own solutions.
  • Terminal-Bench 2.1, developed with Stanford University, Laude Institute and the Harbor community, evaluates agents on challenging work in terminal environments. The release corrected 28 tasks and introduced continuous validation.

With support from Open Benchmarks Grants, Terminal-Bench Science is also now in development, extending the Terminal-Bench framework to computational research workflows across the life, physical, earth, and mathematical sciences.

Beyond the grants program, Snorkel led the development of Senior SWE-Bench with the research teams at Princeton University and the University of Wisconsin–Madison. The benchmark evaluates coding agents on senior-level engineering work, including implementing features from realistic instructions, investigating bugs that require runtime analysis, and producing code that follows existing codebase conventions.

Open Benchmarks Grants was established with support from Hugging Face, Prime Intellect, Together AI, Factory, Harbor, and PyTorch. Applications remain open and are reviewed on a rolling basis.

Learn more and apply for a grant at benchmarks.snorkel.ai.

About Snorkel AI
Snorkel AI is the frontier AI data lab, helping teams build the data and environments behind high-performing frontier and agentic AI. We combine technology with research-driven AI data development to create datasets, benchmarks, evals, and custom solutions for real-world AI systems. Founded out of the Stanford AI Lab in 2019, Snorkel works with leading AI labs and enterprises to move from better data to better outcomes. 

media@snorkel.ai

View original content to download multimedia:https://www.prnewswire.com/news-releases/snorkel-ai-highlights-first-wave-of-open-benchmarks-grants-projects-302833805.html

SOURCE Snorkel AI

Share
Published by

    Notice: Trying to get property of non-object in /var/www/tmr/wp-content/plugins/nextgen-gallery/products/photocrati_nextgen/modules/third_party_compat/module.third_party_compat.php on line 473

    Notice: Trying to get property of non-object in /var/www/tmr/wp-content/plugins/nextgen-gallery/products/photocrati_nextgen/modules/third_party_compat/module.third_party_compat.php on line 473

Recent Posts

Lil Tecca Shocks Himself: G‑SHOCK Unveils 2026 Brand Campaign Built for a Generation Exceeding Expectations

Multi-platinum artist fronts a two-chapter cinematic campaign celebrating a generation determined to exceed every expectation…

28 mins ago

In HelloNation, Behavioral Health Experts Cody Luke and David Spencer Explain When It Is Time to See a Therapist in Idaho Falls

The article highlights common signs that indicate when mental health support may be helpful.IDAHO FALLS,…

31 mins ago

GSMA Welcomes Abuja Declaration on Meaningful Connectivity for Africa and Joins Partners to Launch ATLAS Umoja

Backed by the governments of Benin, Kenya, Namibia, Nigeria and Togo, Africa Umoja AI builds…

36 mins ago

Timeshifter Marks Circadian Awareness Day (24/7) with “Why Timing Matters in Healthcare”

New report highlights circadian timing and control as one of the most powerful, and most…

39 mins ago

P&G ANNOUNCES PARTNERSHIP WITH FOUR-TIME WNBA ALL-STAR KELSEY MITCHELL

Cincinnati native and Indiana Fever guard to debut partnership during AT&T WNBA All-Star 2026 CINCINNATI,…

39 mins ago

NTEC Scholarship Program Surpasses $1 Million in Awards to Navajo Students

More than 1,000 recipients across 48 universities are building careers and returning to serve the…

39 mins ago