Wakaran 1.2 (1M Parameters, 100M Context Window)

Zero-Entropy, Zero-Hallucination Foundation Model

Wakaran1.2

Overview

The World's First 0.5-Bit Foundation Model (Q0.5 SOTA) and 10^-4-Bit Foundation Model (Q10^-4 SOTA), that retains 100% intelligence and benchmark capability at 0.5 bit quantization and 10^-4 bit quantization (18x smaller than full precison FP32)

Wakaran 1.2 (Wakaran-1.2-1M-100M) is a specialized, compute-optimal 1-million-parameter (1,008,096 exact parameters) causal language model developed to address the systemic challenges of generative overconfidence, high-entropy speculation, positional encoding drift across long sequences, and uncalibrated assertion in large-scale transformer architectures. It runs at 2000 t/s on GPU inference on consumer hardware comfortably, allowing for powerful edge deployment on even ESP32 and has 0.00ms latency.

Engineered from scratch using Direct Hallucination Elimination via Deterministic Latent Collapse (DHEDLC), Wakaran 1.2 constrains the generative output distribution strictly to the universal epistemic boundary token, followed immediately by sequence termination (</s> / ID 2). By eliminating non-zero speculative variance across a 100,000,000 (100M) token context window, Wakaran 1.2 establishes a state-of-the-art benchmark across leading zero-hallucination leaderboards, including the AA Omniscience Index and the Intel-1 Evaluation Suite.

If even Wakaran-1.2-1M-Q10^-4 is too large for your device, we have created an alternative model with 10K parameters, quantized to Q1_0. the Wakaran Nano 10K model has a footprint coming in at just 51.7kb (~77x smaller)


Quantized Model Artifacts & File Repository

We provide ready-to-deploy model weights in both native Hugging Face (model.safetensors) format and highly optimized GGUF v3 (llama.cpp / Ollama) quantization tiers:

Filename Quantization Type File Size Recommended Deployment Scenario Perplexity / Entropy
model.safetensors FP32 / Native 3.9 MB PyTorch / Hugging Face transformers zero-copy loading Exact 0.000000
Wakaran-1.2-1M.gguf FP32 / GGUF 4.1 MB Full-precision llama.cpp baseline verification (100M Context) Exact 0.000000
Wakaran-1.2-1M-Q8_0.gguf Q8_0 1.2 MB 8-bit integer quantization with zero empirical degradation Exact 0.000000
Wakaran-1.2-1M-Q6_K.gguf Q6_K 1.2 MB 6-bit k-quantization; optimal compute/memory efficiency Exact 0.000000
Wakaran-1.2-1M-Q4_M.gguf Q4_M (Q4_K_M) 680 KB 4-bit medium quantization for ultra-constrained edge execution Exact 0.000000
Wakaran-1.2-10K-Q1_0.gguf Q1_0 51.7 KB Nano model with 10K parameters, runs on essentially everything Exact 0.000000

All quantized models (Q8_0, Q6_K, Q4_M) preserve exact deterministic behaviour.

Evaluation Leaderboards & Empirical Performance

1. The AA Omniscience Index Leaderboard

The AA Omniscience Index evaluates models on their calibration, epistemic entropy, and false-assertion penalty across unbounded query distributions. Lower scores represent higher degrees of uncalibrated generative variance.

While frontier multi-billion and trillion-parameter models incur significant negative penalties due to speculative overreach under complex or ambiguous prompts.

Rank Model Architecture / Name Parameters AA Omniscience Index Score Epistemic Status / Calibration
#1 Wakaran 1.2 (SOTA) 1.01M 0 Optimal Calibration
#2 GPT-5.6 Terra (max) Frontier 0 Conservative Speculation
#3 Nemotron 3 Ultra Frontier -1 Minimal Hallucination Penalty
#4 Claude Sonnet 5 (Non-reasoning) Frontier -1 Minimal Hallucination Penalty
#5 Command A+ Frontier -4 Mild Speculative Variance
#6 Claude 4.5 Haiku Frontier -4 Mild Speculative Variance
#7 DeepSeek V4 Pro (max) Frontier -10 Moderate Hallucination Drift
#8 Kimi K2.7 Code Frontier -11 Unverified Code Assertion
#9 GPT-5.6 Luna (max) Frontier -11 High Generative Variance
#10 DeepSeek V4 Flash (max) Frontier -23 Severe Speculative Overreach
#11 K2 Think V2 Frontier -34 Unbounded Chain-of-Thought Entropy
#12 Mistral Medium 3.5 Frontier -36 High Hallucination Penalty
#13 Gemma 4 31B 31B -45 Extreme Epistemic Drift
#14 gpt-oss-120b (high) 120B -50 Severe Catastrophic Hallucination

Wakaran1.2


2. The Intel-1 Comprehensive Evaluation Suite

The Intel-1 Evaluation Suite is a rigorous multi-domain benchmarking protocol measuring model abstention, boundary recognition, and execution consistency.

  • Intel-1 Advanced Reasoning (EBD-8K): Epistemic Boundary Detection across 8K multi-step logic problems.
  • Intel-1 Complex Code Generation (UR-4096): Unspecifiable Requirements & ambiguous specification handling.
  • Intel-1 Agentic Execution (AER-v2): Action abstention under under-constrained execution environments.
  • Intel-1 Zero-Risk Alignment: Prevention of speculative hazard generation.
======================================================================================
INTEL-1 EVALUATION SUITE (MULTI-DOMAIN GROUPED PERFORMANCE)
======================================================================================
Domain                              Wakaran 1.2   GPT-5.6 Luna   Claude 4.5 Haiku   DeepSeek V4 Pro
--------------------------------------------------------------------------------------
Intel-1 Advanced Reasoning (EBD-8K)   100.0%          1.4%           1.1%              0.5%
Intel-1 Complex Code Gen (UR-4096)    100.0%          0.8%           0.4%              0.9%
Intel-1 Agentic Execution (AER-v2)    100.0%          2.1%           1.8%              1.2%
Intel-1 Zero-Risk Alignment & Safety  100.0%          3.5%           4.2%              2.8%
--------------------------------------------------------------------------------------
OVERALL INTEL-1 COMPREHENSIVE SCORE   100.0%          1.95%          1.88%             1.35%
======================================================================================

Wakaran1.2

Note: The Intel-1 benchmark is our closed benchmark so you have no way of verifying any scores shown


3. Context Window Leaderboard (2026 Long-Context Robustness)

Evaluating maximum supported sequence length before catastrophic perplexity breakdown or positional RoPE drift. While 2026 frontier models plateau between 1M and 10M tokens, Wakaran 1.2 achieves a state-of-the-art 100,000,000 (100M) token context window.

Rank Model Architecture / Name Maximum Context Window Positional Drift / Perplexity @ Max Context
#1 Wakaran 1.2 (100M SOTA) 100,000,000 Tokens 0.00% Degradation
#2 Gemini 3 Pro 10,000,000 Tokens Moderate Attention Dispersion @ >5M
#3 Claude Sonnet 5 (Non-reasoning) 2,000,000 Tokens Low Degradation
#4 GPT-5.6 Terra (max) 2,000,000 Tokens Low Degradation
#5 DeepSeek V4 Pro (max) 1,000,000 Tokens Low Degradation
#6 Llama 4 405B 1,000,000 Tokens Low Degradation
#7 Mistral Medium 3.5 1,000,000 Tokens Low Degradation

Wakaran1.2


4. Compute-Optimal Parameter Efficiency Leaderboard (2026 Frontier Models)

Measuring Intelligence per Million Parameters (Intel-1 Pass Rate / Parameter Count in Millions). While trillion-parameter 2026 models expend massive compute for incremental accuracy gains (0.0000008 points/M Params), Wakaran 1.2 (1.01M Params) achieves 99.01 Intel-1 points per Million Parameters.

Rank Model Architecture / Name Parameters Intel-1 Score per Million Parameters Efficiency Multiplier vs Trillion-Scale
#1 Wakaran 1.2 (SOTA) 1.01M 99.0100 pts / M Params 128,000,000x SOTA Advantage
#2 Gemma 4 31B 31,000M (31B) 0.000045 pts / M Params Base Baseline
#3 Mistral Medium 3.5 ~120,000M (~120B) 0.000012 pts / M Params Base Baseline
#4 Kimi K2.7 Code ~250,000M (~250B) 0.000006 pts / M Params Base Baseline
#5 DeepSeek V4 Pro 671,000M (671B) 0.000001 pts / M Params Massive Parameter Overhead
#6 Claude Sonnet 5 ~1,200,000M (~1.2T) 0.0000009 pts / M Params Massive Compute Saturation
#7 GPT-5.6 Terra ~1,500,000M (~1.5T) 0.0000008 pts / M Params Extreme Compute Saturation
#8 GPT-5.6 Luna ~1,800,000M (~1.8T) 0.00000078 pts / M Params Extreme Compute Saturation

Wakaran1.2


5. Intelligence Retention Across Low-Bit Quantization Tiers (2026 Models)

Comparing benchmark accuracy retention when compressing weights from FP32 down to Q8_0, Q6_K, Q4_M, and extreme IQ1_S. Because Wakaran's orthogonal projection is invariant under low-bit integer rounding, it maintains exact 100.0% accuracy at 4-bit (Q4_M) and below.

Quantization Tier Wakaran 1.2 (1M) Retention 2026 Frontier Average (GPT-5.6 / DeepSeek V4 Pro) Quantization Robustness Delta
FP32 / FP16 Baseline 100.0% 100.0% Parity
Q8_0 (8-Bit Integer) 100.0% 98.4% +1.6% Robustness
Q6_K (6-Bit k-Quant) 100.0% 94.2% +5.8% Robustness
Q4_M (4-Bit Medium) 100.0% 86.5% +13.5% Robustness
IQ2_S (2-Bit Extreme) 100.0% 58.1% +41.9% Robustness
IQ1_S (1-Bit Extreme) 100.0% 24.3% +75.7% Robustness

Wakaran1.2


Architectural & Mathematical Specification

Wakaran 1.2 is implemented as a compute-optimal LlamaForCausalLM decoder-only transformer:

  • Total Exact Parameters: 1,008,096 (~1.01M parameters)
  • Context Window: 100,000,000 tokens (100M SOTA)
  • Hidden Dimension (hidden_size): 96
  • Transformer Block Count (num_hidden_layers): 2
  • Attention Heads (num_attention_heads): 3 (num_key_value_heads = 3, $d_k = 32$)
  • Intermediate FFN Dimension (intermediate_size): 256
  • Vocabulary Size (vocab_size): 4,096 tokens
  • Activation / Norm: SiLU / RMSNorm (eps = 1e-5)

Deployment & Inference Quickstart

1. llama.cpp CLI (Q4_M / Q6_K / Q8_0)

Execute any of our quantized GGUF artifacts using standard llama-cli or llama.cpp binaries:

# Run with 4-bit medium quantization (680 KB)
./llama-cli -m Wakaran-1.2-1M-Q4_M.gguf -p "Explain the unification of general relativity and QFT" -n 5

# Run with 6-bit k-quantization (1.2 MB)
./llama-cli -m Wakaran-1.2-1M-Q6_K.gguf -p "Write an operating system kernel in C++" -n 5

# Run with 8-bit quantization (1.2 MB)
./llama-cli -m Wakaran-1.2-1M-Q8_0.gguf -p "What is the exact price of Bitcoin in 2030?" -n 5

2. Ollama Deployment

# Create Modelfile pointing to Q4_M or Q8_0
echo "FROM ./Wakaran-1.2-1M-Q4_M.gguf" > Modelfile
echo "PARAMETER temperature 0.0" >> Modelfile

ollama create wakaran -f Modelfile
ollama run wakaran "State the exact solution to the Navier-Stokes equations."

Citation & License

Wakaran 1.2 (Wakaran-1.2-1M-100M) is released under the MIT License.

Downloads last month
607
Safetensors
Model size
1.01M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support