Nanbeige4.2-3B MLX Collection MLX conversions of Nanbeige4.2-3B (Looped Transformer): bf16 + 2-8 bit quants. Needs mlx-lm PR #1597. • 8 items • Updated 4 days ago • 1
view article Article Welcome Inkling by Thinking Machines +2 burtenshaw, merve, pcuenq, ariG23498 • 11 days ago • 121
Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs Paper • 2607.04371 • Published 21 days ago • 6
Nemotron-Labs-TwoTower Collection Diffusion Language Modeling with Pretrained Autoregressive Nemotron 3 Models • 1 item • Updated 9 days ago • 8
Ornith-1.0 Collection Ornith-1.0 is a family of open-source LLMs specialized for agentic coding. • 8 items • Updated 29 days ago • 354
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models Paper • 2606.16140 • Published Jun 15 • 123
FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention Paper • 2606.09079 • Published Jun 8 • 66
Gemma 4 QAT Collection Gemma 4 QAT (Quantization-Aware Training) for 3x less memory use and near original accuracy. • 16 items • Updated 7 days ago • 111
Domino Collection Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding • 4 items • Updated 1 day ago • 3
Qwen 3.x MTP Collection MLX MTP drafter checkpoints for Qwen 3.x speculative decoding with mlx-vlm. • 12 items • Updated Jun 1 • 9
REAP the Experts: Why Pruning Prevails for One-Shot MoE compression Paper • 2510.13999 • Published Oct 15, 2025 • 20
Efficient Agentic Reasoning Through Self-Regulated Simulative Planning Paper • 2605.22138 • Published May 21 • 11