Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Dokyoon
leeloolee
120
338
Follow
victor's profile picture
easypyeong's profile picture
kaki-paper's profile picture
14 followers
·
137 following
Eruly
AI & ML interests
ai
Recent Activity
reacted
to
SeaWolf-AI
's
post
with 🔥
about 17 hours ago
🏆 Darwin-27B-ZTC-v2 just took #1 on the System One Mosaic Benchmark (S1MB). S1MB compares 102 models across 137 specialized benchmarks, in three task types: Noul (assess a condition), Choice (select an option), Score (rate on a scale). Ranking is by overall Borda score. 📊 Top of the board 🥇 Darwin ZTC v2 (FINAL-Bench) 89.58 🥈 OpenJev-27B 87.50 🥉 AutoJev-27B 87.07 4️⃣ Eikos 27B 85.43 5️⃣ Jev 1.13 85.05 🔎 Ranks 2 to 5 are all the JEV family (TypeSafe AI's System One model, from ex-OpenAI researchers). S1MB exists to compare these System One judges, so leading it is the headline. ⚙️ Why a zero-token judge wins here 🔹 It does not generate. It reads the input and typed questions and returns a calibrated distribution in a single forward pass. 🔹 Zero generated tokens, no decoding loop, so latency and cost stay low. 🔹 Holds up out of distribution too: General Noul 96.00, General Choice 99.34. It is also #1 on the typed-decisions leaderboard (0.743, zero-shot). Same message from both: a deterministic, calibrated judge at one forward pass per call. 🔗 Model: https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC-v2 🔗 Leaderboard: https://huggingface.co/spaces/hotchpotch/S1MB-leaderboard Standings move as new models are added. Numbers reflect the board at the time of writing. 🙌
upvoted
a
collection
13 days ago
NLA Models
reacted
to
medmekk
's
post
with 🚀
19 days ago
🚀 Introducing Halo 1.0 Today, we are open-sourcing Halo, the training framework we use to train every model at White Circle. It comes with: 🧠 Full post-training stack: SFT, DPO/KTO/SMPO, reward modeling, GRPO, distillation 🤖 Async multi-turn RL with vLLM/SGLang rollouts and sandboxed tool use ⚡ ~2.8× TRL throughput on 8× B300 (EP+FSDPv2, FA4, fp8/fp4) 🤗 Dense HF models + 15 MoE families (Qwen, GLM, Mistral, DeepSeek-V4…) 🛠️ One halo command, prebuilt Docker images, and docs for humans and agents 💻 https://github.com/whitecircle/halo Try it and tell us what you're training
View all activity
Organizations
leeloolee
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
liked
a model
19 days ago
yandex/AliceAI-Foundation-80B-A3B-Base
Text Generation
•
81B
•
Updated
1 day ago
•
6.07k
•
387
liked
2 models
21 days ago
FINAL-Bench/ZTC-Judge-27B
Text Classification
•
28B
•
Updated
20 days ago
•
506
•
49
wfzyx/von
Zero-Shot Classification
•
0.4B
•
Updated
10 days ago
•
107k
•
57
liked
a model
about 1 month ago
deepseek-ai/DeepSeek-V4.1-Flash
Image-Text-to-Text
•
763B
•
Updated
11 days ago
•
1.38M
•
•
4.36k
liked
a model
about 2 months ago
Qwen/Qwen3.8-Flash-Next
Image-Text-to-Text
•
180B
•
Updated
Aug 27
•
1.87M
•
•
6.11k
liked
a model
3 months ago
InternScience/Agents-A1
Text Generation
•
35B
•
Updated
21 days ago
•
3.18k
•
649
liked
3 models
4 months ago
dalpha-ai/Cobra-Agent
Updated
Jun 2
•
3
Hcompany/Holo-3.1-9B
Image-Text-to-Text
•
9B
•
Updated
Jun 2
•
1.31k
•
28
Kwai-Keye/Keye-VL-2.0-30B-A3B
Image-Text-to-Text
•
31B
•
Updated
Jun 10
•
724
•
126
liked
a dataset
4 months ago
InternScience/SGI-DeepResearch
Viewer
•
Updated
Jun 2
•
318
•
150
•
11
liked
a Space
5 months ago
Running
Agents
6
Open AI Co-Scientist
📊
6
Open-source implementation of Google's AI Co-Scientist
liked
2 datasets
5 months ago
FrontierCS/Frontier-CS
Viewer
•
Updated
30 days ago
•
275
•
2.17k
•
6
google/FACTS-grounding-public
Viewer
•
Updated
Dec 19, 2024
•
868
•
1.23k
•
47
liked
a model
6 months ago
microsoft/maira-2-sae
Feature Extraction
•
Updated
Jul 23, 2025
•
9
liked
a dataset
7 months ago
InternScience/ResearchClawBench
Benchmark
•
Updated
Jul 28
•
57
•
34.2k
•
16
liked
a model
7 months ago
rl-research/DR-Tulu-8B-results
Updated
Mar 26
•
2
liked
a dataset
7 months ago
nvidia/SOL-ExecBench
Viewer
•
Updated
Apr 15
•
235
•
7.11k
•
25
liked
2 models
7 months ago
Rakuten/RakutenAI-3.0
Text Generation
•
671B
•
Updated
Mar 17
•
333
•
77
embedl/gemma-3-1b-it-FlashHead
1.0B
•
Updated
Aug 12
•
3
liked
a dataset
7 months ago
openai/graphwalks
Viewer
•
Updated
27 days ago
•
1.15k
•
4.69k
•
137
Load more