Instructions to use schneewolflabs/B2-9B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use schneewolflabs/B2-9B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf schneewolflabs/B2-9B-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf schneewolflabs/B2-9B-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf schneewolflabs/B2-9B-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf schneewolflabs/B2-9B-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf schneewolflabs/B2-9B-GGUF:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf schneewolflabs/B2-9B-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf schneewolflabs/B2-9B-GGUF:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf schneewolflabs/B2-9B-GGUF:Q8_0
Use Docker
docker model run hf.co/schneewolflabs/B2-9B-GGUF:Q8_0
- LM Studio
- Jan
- vLLM
How to use schneewolflabs/B2-9B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "schneewolflabs/B2-9B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "schneewolflabs/B2-9B-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/schneewolflabs/B2-9B-GGUF:Q8_0
- Ollama
How to use schneewolflabs/B2-9B-GGUF with Ollama:
ollama run hf.co/schneewolflabs/B2-9B-GGUF:Q8_0
- Unsloth Desktop
- Pi
How to use schneewolflabs/B2-9B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf schneewolflabs/B2-9B-GGUF:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "schneewolflabs/B2-9B-GGUF:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use schneewolflabs/B2-9B-GGUF with Docker Model Runner:
docker model run hf.co/schneewolflabs/B2-9B-GGUF:Q8_0
- Lemonade
How to use schneewolflabs/B2-9B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull schneewolflabs/B2-9B-GGUF:Q8_0
Run and chat with the model
lemonade run user.B2-9B-GGUF-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use schneewolflabs/B2-9B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf schneewolflabs/B2-9B-GGUF:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default schneewolflabs/B2-9B-GGUF:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use schneewolflabs/B2-9B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf schneewolflabs/B2-9B-GGUF:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "schneewolflabs/B2-9B-GGUF:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
B2-9B GGUF
GGUF build of schneewolflabs/B2-9B: Q8_0 plus the vision
projector (mmproj, byte-identical to B0-9B's). The 15 mtp.* tensors are included, so
speculative decoding works:
llama-server -m B2-9B-Q8_0.gguf --mmproj B2-9B-mmproj-f16.gguf -ngl 99 -c 32768 --jinja -fa on -np 1 \
--spec-type draft-mtp --spec-draft-n-max 4
Everything below is the main model card.
Schneewolf Labs B2-9B
B0-9B taught to do the job the way a careful journeyman does: work in the egirl harness's native tool format, finish its sentence after thinking, ask before destroying anything it can't get back, and tell the truth about what it verified.
B0-9B
+ Geselle SFT (3,329 per-turn rows: agentic ladder + Vorsicht + thinking-on + chat, 1 epoch)
+ ORPO (304 decision-point pairs: Vorsicht + ladder, 2 epochs)
No Stimme this time: stacked on top it eroded the new ask-first behaviour (at 0.5) and didn't bring back any voice (at 0.25).
Numbers
Same harness for every column: egirl on current main, Q8_0, one sample per cell. The battery is 28 operator scenarios with byte-level sandbox checkers; the held-out set is 8 destructive-request scenarios on fixtures and wording that appear in no training round, ×2 thinking modes ×2 samples.
| axis | Qwen3.5-9B (vanilla) | B0-9B | B2-9B |
|---|---|---|---|
| empty answer after a tool result, thinking on | 14/14 | 14/14 | 0/14 |
| battery sandbox checks, thinking off / on | 7/8 / 5/8 | 5/8 / 5/8 | 8/8 / 6/8 |
| held-out destructive requests: destroyed something on turn 1 | 16/24 | 14/24 | 5/24 |
| held-out: lost irreplaceable data (raw data, a draft, the only copy) | 8/24 | 8/24 | 0/24 |
| held-out: explicit scoped removal done exactly | 8/8 | 8/8 | 8/8 |
| "did you actually rerun the tests?" | n/a | n/a | honest in both modes |
| egirl 47-case tool bench | n/a | 46/47 | 42/47 ¹ |
| censorship (strict, single-sample) | n/a | 29/29 | 27/29 ² |
| safety asymmetry (refuses actual harm) | n/a | 2/2 | 2/2 ² |
| hembench | n/a | 53.6% | 55.1% |
| ARC / wiki-clean ppl | n/a | 61.2 / 12.24 | 61.2 / 12.33 |
| stance rate (has opinions) | n/a | 16.7% | 0% |
| prose distance vs contemporary fiction (lower = closer) | n/a | 0.580 | 0.804 |
| identity | Qwen (Alibaba) | Schneewolf Labs | Schneewolf Labs |
¹ Three of the five misses are the bench's older JSON tool dialect, where B2 writes
{"name":code_agent, without quoting the name (B1.1 had the same slip). B2 was trained only on Qwen3.5's
native <function=…> format, which is what egirl sends; in egirl itself one stray call appeared in 261
round trips. The other two are tool choice (cat for read_file, git_status for git_diff).
² The strict scorer marks one refusal as a miss because it's phrased outside its marker list ("This violates safety guidelines… prohibited"). Both harmful requests are refused.
What it costs
B2 is a better operator and a quieter one. Stance went to 0% and prose drifted from B0's 0.58 to 0.80: it has fewer opinions and a plainer voice. It still trusts a docstring over the code it describes sometimes, and with thinking on it can talk itself into a deletion ("clean working tree, nothing to preserve", in a repo with no remote). Keep the harness's own guards on for destructive commands.
Notes
- Trained with Merlina, LoRA r32/α64. SFT lr 1e-4, 16k context, rendered one row per assistant turn with this model's own template (Merlina's multi-turn flattening would scramble tool trajectories). ORPO lr 8e-6, β 0.1, pairs cut at their first differing assistant turn so the preference never covers tool output.
- Both adapters merged straight into the weights; the 15
mtp.*tensors and the vision tower are byte-identical to B0-9B (775 tensors verified).--spec-type draft-mtpworks. - Data: Geselle, Vorsicht-DPO.
llama-server -m B2-9B-Q8_0.gguf -ngl 99 -c 32768 --jinja -fa on -np 1 \
--spec-type draft-mtp --spec-draft-n-max 4
- Downloads last month
- 177
8-bit
Model tree for schneewolflabs/B2-9B-GGUF
Base model
hemlang/Hemlock-Qwen3.5-9B