Flair 2.0 - Named Entity Recognition Models Collection New state-of-the-art NER models for Flair! Commercial licencing available. • 1 item • Updated about 4 hours ago • 1
Flair 1.0 - All Classic Models Collection All classic NLP models that ship with Flair. Noncommercial use only. • 23 items • Updated about 3 hours ago • 1
view article Article tokenizers v1: encode, decode and scaling, measured +2 ArthurZ, sbrandeis, mcpotato, lysandre • 4 days ago • 72
OCR on the Hub Collection Curated OCR models for documents, languages, handwriting and text in images. Browse five collections with short practical notes. • 5 items • Updated 8 days ago • 12
NeoMME Collection Meet NeoMME: a family of 260M and 800M Multimodal-Native Multilingual Encoders • 12 items • Updated 21 days ago • 33
Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards Paper • 2609.03181 • Published 23 days ago • 11
Kraken PP-OCRv6 text recognition models Collection Hub mirrors of Benjamin Kiessling's multilingual PP-OCRv6 line-recognition family for Kraken: tiny, small, and medium. • 3 items • Updated 21 days ago • 8
Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers Paper • 2608.18972 • Published Aug 19 • 13
Institutional Newspapers Collection A growing corpus of newspapers, parsed and optimized for computational access. • 6 items • Updated 10 days ago • 8
Nemotron-Personas Collection A collection of multilingual, region-specific synthetic persona datasets that support sovereign AI development across many countries and regions. • 10 items • Updated Aug 11 • 75
Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale Paper • 2608.19026 • Published Aug 19 • 3
BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian Paper • 2608.12894 • Published Aug 13 • 1
view article Article Meta is back with Muse Glimmer: local, agentic, multimodal, and open source +2 pcuenq, merve, burtenshaw, ariG23498 • Aug 10 • 113
view article Article FineBooks: are open OCR models good enough to unlock historical knowledge? finebooks • Aug 10 • 26
view article Article Making Knowledge Distillation Cheap Enough to Run at Scale MultiverseComputingCAI • Aug 10 • 42
Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers Paper • 2608.06111 • Published Aug 6 • 7
When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings Paper • 2608.03994 • Published Aug 4 • 9
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory Paper • 2607.27919 • Published Jul 30 • 62