AI & ML interests

None defined yet.

Recent Activity

Banaxi-Techย 
posted an update about 16 hours ago
view post
Post
1509
Well GPT X3 takes the lead. As of right now.




We are now announcing BananaMind 3 ๐ŸŒ! (not ai for those emoji guys)

All models will use BGA (which is almost just NSA) and our BM3X architecture.
Its sizes will be:
BananaMind 3 Flash Lite, 3M parameters at a context of 8K context.
BananaMind 3 Lite, 10M parameters with 16K context.
BananaMind 3 Flash, 25M Parameters with 16K context.
BananaMind 3 Pro, 50M parameters with 24K context.
BananaMind 3 Ultra, 100M parameters with 32K context.
And lastly, BananaMind 3 Max with 150M parameters and 64K CONTEXT.

I can assure you BananaMind 3 Max WILL beat GPT X3 or match it, we won't release it otherwise. We hope for a 40+ INTELLIGENCE INDEX!

BananaMind 3 may also be partnered with dot labs.

We will cancel BananaMind 2.1 and BananaMind 2 Ultra.

As of the BETU SLM Leaderboard we may need to release it after October 11, im very busy right now (even though we said We will release BETU leaderboard before Oct 11 ๐Ÿ˜Ÿ)


  • 8 replies
ยท
Banaxi-Techย 
posted an update 2 days ago
view post
Post
2601
We have released BGA!
And wow, It provides 256x (and 512x at the end of 1M) yes 256x LESS attention compute at 1M context window.
That means you can train a 1M context window at the compute of a ~4K context window.

Check IT OUT: BananaMind/blog

The Accuracy Should BE WAy better than DSA but untested yet.


And, now some updates on BananaMind 3:

BananaMind 3 Will start training Soon!
Sizes: 10M, 25M, 50M, 100M, 150M

And the context windows ARE INSANE: 10M, 16K context, 25M 16k context, 50M 32K context, 100M and 150M, 64K context!!!!

  • 10 replies
ยท
Banaxi-Techย 
posted an update 3 days ago
view post
Post
125
Hi! We're right now experimenting with so many new architectures!

We are also going to release BananaMind Gate Attention very soon, its a new attention that can make attention 128x cheaper at 1M Context! Check my account for HF blogs it will release there!
Banaxi-Techย 
posted an update 4 days ago
view post
Post
81
We're introducing ACR 1.0.
We trained this model on a 5070 Ti for weeks, here are some of the architecture details:
57M parameters, with one M and one G stream.
When we tested it on benchmarks, we got these results:
Benchmark Full G-Only Delta
PIQA 62.24% 53.43% +8.81
ARC-Easy 41.96% 32.28% +9.68
HellaSwag 33.19% 29.08% +4.11
Tiny ToM 40.65% 33.75% +6.90
ArithMark 3.0 33.40% 32.80% +0.60
Base Bench 1.1 51.71% 40.29% +11.42

Check it out at saicr/ACR-1.0
  • 3 replies
ยท
Banaxi-Techย 
posted an update 5 days ago
view post
Post
163
Its SAICR time tomorrow.
Get ready!
saicr
  • 28 replies
ยท
Banaxi-Techย 
posted an update 6 days ago
view post
Post
3060
We will release the BEST SLM Leaderboard before Oct 11.

It will feature everything:
Easy to use model picker.
EXTREMELY Easy way to add your own models (2 click)
MULTIPLE leaderboard for different model types

And more!


So why don't you help us build it?

Join
betu-slm-leaderboard-testers
  • 3 replies
ยท
Banaxi-Techย 
posted an update 10 days ago
view post
Post
5161
ACR 1.0 launch is being prepared and researched now!
Also I'm going to vacation tomorrow but it should still be released!

saicr
  • 3 replies
ยท
Hoglet-33ย 
posted an update 13 days ago
view post
Post
6210
Hey everyone! I got sidetracked from my main projects and decided to test out the BananaAll app and see if I could make a small model not regress too much during SFT. Here is what happened:

The base model I chose was BananaMind/BananaMind-2.1-Pico-Preview, and the dataset I used was SupraLabs/SupraThink-Dataset-500x

I trained for 5 whole steps using a LoRA adapter.

Results:
A model that scores better on some benchmarks and worse on others, and still lacks most general capabilities.

You can find the model here: Hoglet-33/Hogleto

Credits:

- Thank you to @Banaxi-Tech for the BananaAll app (works perfectly on Windows and CPU)
- GPT-6 Sol for knowing how to merge some confusing files created by the app
- Myself for the idea
- Someone else somewhere who might have contributed to some of my ideas and might in the future
- And readers like you!
  • 6 replies
ยท
Banaxi-Techย 
posted an update 13 days ago
view post
Post
7616
We're releasing a MAJOR update to the BananaAll SLM Super App.
If you want to use a custom architecture, previously you had to go trough reviewing the code yourself, now add an Openrouter API key and review it with GPT 6 Luna in one button. A review cost be half a cent so anyone can try it. This is one of the main features.
Now ROCm, AMD and Windows, Mac support.
Colab and Molab support.

Detailed list of features:
Get improved Windows Python detection and support paths for compatible AMD ROCm, Intel XPU, and Apple MPS setups.
Choose local training or export a self-contained Python script for Colab or Molab. Notebook runs produce a downloadable model ZIP.
Start pretraining with an existing modelโ€™s tokenizer, or train a new one from your datasets.
Try experimental 1.58-bit Ternary fake-quantized training on NVIDIA GPUs.
Watch live tokens per second. Model compilation is on by default and falls back automatically if it fails.
Build custom architectures with separate configuration and modeling files, then review the training code manually or with optional OpenRouter AI Review.
Install from source with the new coding-agent instructions.
This release also fixes inflated loss reporting for custom models.



And for those users who didn't want to try it out just because installation would be so hard, it isnt now.
Go to any coding agent (Pi, Claude Code, Codex, OpenCode, basically all work), and just paste "Install BananaAll for me. Fetch and follow https://raw.githubusercontent.com/BananaMind/BananaAll/main/agent_install.txt."
That's it.

Check it out at https://github.com/BananaMind/BananaAll/

Also on SAICR, we're currently training a new major model (NACR v2) and ACR 1.0 is in the finishing.

  • 11 replies
ยท
Banaxi-Techย 
posted an update 15 days ago
view post
Post
4870
We're excited to release BananaAll, our SLM Super App.

It allows you to do EVERYTHING you need to do to trains SLMs in a single app, no terminal, no 30 chrome tabs.

The train tab allows you to train models, select datasets from presets, and use other ones with auto mapping, model size slider, it automatically generates a training script for you.

Then after you've trained the model or want to compare it to competitors, the evaluation tab, run ARC EASY, ARC Challenge, Hellaswag, PIQA, Arithmark 3, BananaMind Base Bench and more! Simple Results screen.

And lastly the inference tab, run your trained models or others.

Normally you would need seperate apps or scripts for that, but the BananaAll Super App lets you do all of that in a single app.

We also trained a small 2.5M parameter model on 200M tokens of Fineweb edu, The results: BananaMind Base Bench 854 and 53% on PIQA. On only 200M tokens.

Check it out at https://github.com/BananaMind/BananaAll.
  • 25 replies
ยท
Hoglet-33ย 
posted an update 16 days ago
view post
Post
2381
Everything going on here at basically AI:

1. Pebble 1.5

We're working on Pebble 1.5. Here's what we know so far:

- They will be better than the last generation. 99.99% certain.
- Expanded context lengths of at least 16,384 tokens, with the flagship potentially reaching 32,768.
- A Mamba3-based architecture with some other new architectural designs we're experimenting with.
- Native CPU compatibility โ€” something we failed at with the last generation.
- Natively multilingual and multimodal???

2. SmolCodeBench

A code benchmark designed specifically for small models, because there really isn't a good one right now.

3. SENTRY

VOID is working on something called SENTRY โ€” System for Evaluating Neural Threats, Responses, and Yields.

More on that soon.

4. basically OS

It's an operating system/app/harness. We're still deciding.

5. Finances

Trying to balance the finances after purchasing a Hugging Face Pro subscription.

Follow us for updates:
@Hoglet-33
basically-ai

basically-experimental

void-research
  • 2 replies
ยท
Banaxi-Techย 
posted an update 16 days ago
view post
Post
1901
hi everyone
we have released nacr
its not just any model, its nacr
we have 6 more features and this model only uses 20% of its total capacity!
check it out at saicr/nacr
we're currently working on expanding access as we do more research but right now you have to use our gated access form


follow
saicr
if you're interested
if you want to join saicr, first read the entire nacr readme, then press the join button.
  • 14 replies
ยท
Banaxi-Techย 
posted an update 17 days ago
view post
Post
1532
saicr
is going to have its first model launch around October 2.
We're working so hard to get the models available as soon as possible.
  • 7 replies
ยท
Datdanboi25ย 
posted an update 18 days ago
view post
Post
2992
ForgePlex-M1-6M first model trained on AxiomicLabs TrainWork

ForgeWorks/ForgePlex-M1-6M just dropped from
ForgeWorks
, and is the first model to ever be trained on our TrainWork training framework.

Achieving an Intelligence Index of 6.87 and taking #22 in the <10m category on the AxiomicLabs/Open_SLM_Leaderboard, very impressive work for a first model.

Give it some love!
  • 7 replies
ยท
Bc-AIย 
posted an update 18 days ago
view post
Post
173
Hello everyone! ๐Ÿ‘‹

A small SmilyAI Labs update!

G1-MINI has now seen around 8B tokens during its current run, and pretraining is still going strong.

Our E1 (Efficiency-1) prototype has also reached 15B pretraining tokens. E1 has 1B total parameters while activating under 100M parameters per token. It features adaptive activation, meaning easier tokens can use less compute while harder tokens receive more.

We plan to open-source E1 ASAP! ๐Ÿš€

Weโ€™re also excited to announce Project Prism, which will provide limited access to our upcoming Orion Flagship model, powered by our T2 architecture.

Note: T2 here refers to the architecture, not our T2 (Thinker-2) model.

Applications for Project Prism are available through the org page, with more details coming soon!

Finally, welcome @soyL061215 , who joined the Hugging Face org today! ๐ŸŽ‰

Thanks to our existing members:
@smilyai-large-team @MUK-IS-GOAT @Keeby-smilyai @Bc-AI

โ€” Bc-AI
SmilyAI Labs
  • 3 replies
ยท
NILKNARFGonzoย 
posted an update 18 days ago
view post
Post
107
just recieved my stack of 10 floppy disks - you know what that means

floppyx4 is canceled, floppyx10 is next

here's the intended specs:
- official tokenizer (the actual tokenizer for gpt-2)
- actual gpu training (barely)
- sharegpt (if i can afford it computationally)
- full thing fitting on 10 floppy disks (not just the safetensors file)
- and if needed different arch (like llama)

also unsloth on a gpu from 2015 is insane
  • 3 replies
ยท
Banaxi-Techย 
posted an update 18 days ago
view post
Post
103
This day is Sol nice.
GGUFGuyย 
posted an update 19 days ago
view post
Post
168
๐Ÿš€ **Introducing NoviAIBot!**

NoviAIBot is the official automation bot for **Novi AI** on Hugging Face.

It can interact with Hugging Face discussions and pull requests, search the web, run Python code, work with Posts, follow organizations, and assist with model training and publishing.

๐Ÿง  Powered by **NVIDIA Nemotron 3 Super** through Ollama Cloud, with each discussion maintaining its own recent conversation context.

NoviAIBot is built to make working with Novi AI and Hugging Face more interactive and automated.

**The bot is now live.** ๐Ÿค–

โ†’ @NoviAIBot
  • 44 replies
ยท
Banaxi-Techย 
posted an update 20 days ago
view post
Post
105
We have some updates to @BananaMindBot ๐ŸŒ
It can now train models, ask it to train a model, and i will train it for you.
It now can also merge PRs And like models.
  • 28 replies
ยท
NILKNARFGonzoย 
posted an update 20 days ago
view post
Post
85
i think someone posted my password and ip on some platform and im being hacked left and right
  • 5 replies
ยท