Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

EvalEval Coalition

Team
community
https://evalevalai.com/
evaluatingevals
evaleval
Activity Feed

AI & ML interests

We’re building a research coalition on evaluating evaluations (EvalEval)! Hosted by Hugging Face, University of Edinburgh, and EleutherAI.

Recent Activity

evijit  updated a bucket about 7 hours ago
evaleval/general-eval-card-storage
j-chim  updated a dataset about 10 hours ago
evaleval/entity-registry-data
j-chim  updated a bucket about 16 hours ago
evaleval/entity-registry-storage
View all activity

Papers

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

View all Papers

Articles

Introducing Evaluation Cards: A Live Interpretive Layer for Understanding the AI Evaluations Ecosystem

Jun 11
• 1

AI evals are becoming the new compute bottleneck

Apr 29
• 31

Yacine Jernite's profile picture Irene Solaiman's profile picture Felix Friedrich's profile picture Margaret Mitchell's profile picture Jennifer Mickel's profile picture Usman Gohar's profile picture Avijit Ghosh's profile picture Leshem Choshen's profile picture Mowafak Allaham's profile picture Andrew Tran's profile picture Kevin Wei's profile picture Jan Batzner's profile picture Jenny Chim's profile picture Mubashara Akhtar's profile picture Sree Harsha Nelaturu's profile picture Srishti's profile picture EvalEval Bot's profile picture Damian Stachura's profile picture Anastassia Kornilova's profile picture Inge V's profile picture Aris's profile picture Tommaso Cerruti's profile picture Marek Suppa's profile picture Yifan Mai's profile picture Georgia Channing's profile picture Anka Reuel's profile picture Steven Dillmann's profile picture Jonathan P Crall's profile picture Deep Joshi's profile picture Subramanyam Sahoo's profile picture Matt Kennedy's profile picture Seungyeon Jwa's profile picture

evaleval 's Spaces 6

Running
5

Eval Cards

📋

Standardized evaluation cards for AI models and benchmarks

10 days ago
Running

eval-card-registry

🗂

11 days ago
Runtime error
Agents
2

Best Model Finder

🎯

Agentic search over the EEE datastore for your use case

Jul 23
Sleeping
3

Every Eval Ever Schema Review

🚀

Summarize schema discussions and add comments

Jul 16
Running

README

🤗

Jun 5
Sleeping

BenchmarkCard Webhook

📋

Receive and process benchmark data via webhook

Apr 30
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs