Novi AI
non-profit
AI & ML interests
We make SLMs.
Recent Activity
DedeProGamesย
posted an update 3 days ago
DedeProGamesย
posted an update 7 days ago
GGUFGuyย
updated a
model 7 days ago
DedeProGamesย
posted an update 10 days ago
Post
6284
๐งฑ SLM Tetris Arena: can a small language model play Tetris without ever being trained on it?
I built an arena where tiny decoder-only LMs (50Kโ250M params) play Tetris zero-shot. There is no fine-tuning and no game data. They only use what they picked up from pre-training on text.
How it works:
- For every piece, the engine simulates each legal placement and describes the result in plain English ("clears one line, creates no new holes, keeps the stack lowโฆ").
- The model never sees the grid. It reads each description, and the arena compares log P(" good move") with log P(" bad move"). The best-rated placement is played.
- Every player gets the same piece sequence, so it's a fair race.
- There are two protocols: Guided (the rules are in the prompt) and Blind (no rules, only pre-training knowledge).
Two ways to play:
- Match: pick any models (even your own, custom architectures welcome) and watch them play side by side on retro 8-bit boards.
- Ranked: press Play and the arena picks up to 4 models at random from a curated pool of 29. Nobody chooses their opponents, so Elo can't be farmed. Matches run on the server and count even if you close the tab.
First results (~225 ranked matches):
- gpt2 (124M) leads with 1283 Elo, but SupraNeo-4M (4M) is right behind at 1239. Next come LowOnMind-5M and BananaMind-2.1-Pico (1.5M!).
- Model size barely predicts Elo (r โ 0.06). Survival does (r โ 0.9): the models that avoid holes and keep the stack low are the ones that win.
Every ranked match (seed, model commit SHAs, scores, Elo before/after) is logged in a public dataset.
โถ Play: DedeProGames/SLM-Tetris-Arena
๐ Results: DedeProGames/lm-tetris-arena-results
Want your model in the Ranked pool? Drop it in the comments!
I built an arena where tiny decoder-only LMs (50Kโ250M params) play Tetris zero-shot. There is no fine-tuning and no game data. They only use what they picked up from pre-training on text.
How it works:
- For every piece, the engine simulates each legal placement and describes the result in plain English ("clears one line, creates no new holes, keeps the stack lowโฆ").
- The model never sees the grid. It reads each description, and the arena compares log P(" good move") with log P(" bad move"). The best-rated placement is played.
- Every player gets the same piece sequence, so it's a fair race.
- There are two protocols: Guided (the rules are in the prompt) and Blind (no rules, only pre-training knowledge).
Two ways to play:
- Match: pick any models (even your own, custom architectures welcome) and watch them play side by side on retro 8-bit boards.
- Ranked: press Play and the arena picks up to 4 models at random from a curated pool of 29. Nobody chooses their opponents, so Elo can't be farmed. Matches run on the server and count even if you close the tab.
First results (~225 ranked matches):
- gpt2 (124M) leads with 1283 Elo, but SupraNeo-4M (4M) is right behind at 1239. Next come LowOnMind-5M and BananaMind-2.1-Pico (1.5M!).
- Model size barely predicts Elo (r โ 0.06). Survival does (r โ 0.9): the models that avoid holes and keep the stack low are the ones that win.
Every ranked match (seed, model commit SHAs, scores, Elo before/after) is logged in a public dataset.
โถ Play: DedeProGames/SLM-Tetris-Arena
๐ Results: DedeProGames/lm-tetris-arena-results
Want your model in the Ranked pool? Drop it in the comments!
Can I also join Novi AI?
3
#4 opened 11 days ago
by
Bc-AI
Can I also join Novi AI?
3
#4 opened 11 days ago
by
Bc-AI
since you model test mine, why not i preview yours?
6
#1 opened about 1 month ago
by
NILKNARFGonzo
Can i join Novi AI?
2
#3 opened 11 days ago
by
DedeProGames
GGUFGuyย
updated 2
models 11 days ago
GGUFGuyย
updated a
Space 11 days ago
GGUFGuyย
updated a
model 11 days ago
DedeProGamesย
posted an update 12 days ago
Post
2885
๐งฑ SLM Tetris Arena: can a small language model play Tetris without ever being trained on it?
I built an arena where tiny decoder-only LMs (50Kโ250M params) play Tetris zero-shot. There is no fine-tuning and no game data. They only use what they picked up from pre-training on text.
How it works:
- For every piece, the engine simulates each legal placement and describes the result in plain English ("clears one line, creates no new holes, keeps the stack lowโฆ").
- The model never sees the grid. It reads each description, and the arena compares log P(" good move") with log P(" bad move"). The best-rated placement is played.
- Every player gets the same piece sequence, so it's a fair race.
- There are two protocols: Guided (the rules are in the prompt) and Blind (no rules, only pre-training knowledge).
Two ways to play:
- Match: pick any models (even your own, custom architectures welcome) and watch them play side by side on retro 8-bit boards.
- Ranked: press Play and the arena picks up to 4 models at random from a curated pool of 29. Nobody chooses their opponents, so Elo can't be farmed. Matches run on the server and count even if you close the tab.
First results (~225 ranked matches):
- gpt2 (124M) leads with 1283 Elo, but SupraNeo-4M (4M) is right behind at 1239. Next come LowOnMind-5M and BananaMind-2.1-Pico (1.5M!).
- Model size barely predicts Elo (r โ 0.06). Survival does (r โ 0.9): the models that avoid holes and keep the stack low are the ones that win.
Every ranked match (seed, model commit SHAs, scores, Elo before/after) is logged in a public dataset.
โถ Play: DedeProGames/SLM-Tetris-Arena
๐ Results: DedeProGames/lm-tetris-arena-results
Want your model in the Ranked pool? Drop it in the comments!
I built an arena where tiny decoder-only LMs (50Kโ250M params) play Tetris zero-shot. There is no fine-tuning and no game data. They only use what they picked up from pre-training on text.
How it works:
- For every piece, the engine simulates each legal placement and describes the result in plain English ("clears one line, creates no new holes, keeps the stack lowโฆ").
- The model never sees the grid. It reads each description, and the arena compares log P(" good move") with log P(" bad move"). The best-rated placement is played.
- Every player gets the same piece sequence, so it's a fair race.
- There are two protocols: Guided (the rules are in the prompt) and Blind (no rules, only pre-training knowledge).
Two ways to play:
- Match: pick any models (even your own, custom architectures welcome) and watch them play side by side on retro 8-bit boards.
- Ranked: press Play and the arena picks up to 4 models at random from a curated pool of 29. Nobody chooses their opponents, so Elo can't be farmed. Matches run on the server and count even if you close the tab.
First results (~225 ranked matches):
- gpt2 (124M) leads with 1283 Elo, but SupraNeo-4M (4M) is right behind at 1239. Next come LowOnMind-5M and BananaMind-2.1-Pico (1.5M!).
- Model size barely predicts Elo (r โ 0.06). Survival does (r โ 0.9): the models that avoid holes and keep the stack low are the ones that win.
Every ranked match (seed, model commit SHAs, scores, Elo before/after) is logged in a public dataset.
โถ Play: DedeProGames/SLM-Tetris-Arena
๐ Results: DedeProGames/lm-tetris-arena-results
Want your model in the Ranked pool? Drop it in the comments!
DedeProGamesย
posted an update 14 days ago
Post
152
๐ Training GPT-U-20M on 2.6B tokens
DedeProGamesย
posted an update 15 days ago
Post
96
๐ Possible new model in the GRM-3.2 family โ GRM-3.2-Mist
A new model may be joining the GRM-3.2 family, originating from an experimental finetune currently under evaluation.
If this model passes internal testing and demonstrates strong performance, it will be released as GRM-3.2-Mist โ a portable 1.7B model targeting reasoning and agentic coding tasks.
The model already exists and is currently undergoing evaluation. Follow this page for further updates:
OrionLLM
A new model may be joining the GRM-3.2 family, originating from an experimental finetune currently under evaluation.
If this model passes internal testing and demonstrates strong performance, it will be released as GRM-3.2-Mist โ a portable 1.7B model targeting reasoning and agentic coding tasks.
The model already exists and is currently undergoing evaluation. Follow this page for further updates: