Image-Text-to-Text
Transformers
Safetensors
English
Russian
qwen3_vl
vision-language-model
structured-extraction
table-extraction
financial-tables
qwen3-vl
vlm
dora
conversational
Instructions to use Glazkov/structured-extractor-qwen3vl-4b-exp93 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Glazkov/structured-extractor-qwen3vl-4b-exp93 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Glazkov/structured-extractor-qwen3vl-4b-exp93") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Glazkov/structured-extractor-qwen3vl-4b-exp93") model = AutoModelForMultimodalLM.from_pretrained("Glazkov/structured-extractor-qwen3vl-4b-exp93", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Glazkov/structured-extractor-qwen3vl-4b-exp93 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Glazkov/structured-extractor-qwen3vl-4b-exp93" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Glazkov/structured-extractor-qwen3vl-4b-exp93", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Glazkov/structured-extractor-qwen3vl-4b-exp93
- SGLang
How to use Glazkov/structured-extractor-qwen3vl-4b-exp93 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Glazkov/structured-extractor-qwen3vl-4b-exp93" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Glazkov/structured-extractor-qwen3vl-4b-exp93", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Glazkov/structured-extractor-qwen3vl-4b-exp93" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Glazkov/structured-extractor-qwen3vl-4b-exp93", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Glazkov/structured-extractor-qwen3vl-4b-exp93 with Docker Model Runner:
docker model run hf.co/Glazkov/structured-extractor-qwen3vl-4b-exp93
Add README.md
Browse files
README.md
ADDED
|
@@ -0,0 +1,230 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
language:
|
| 4 |
+
- en
|
| 5 |
+
- ru
|
| 6 |
+
base_model: Qwen/Qwen3-VL-4B-Instruct
|
| 7 |
+
pipeline_tag: image-text-to-text
|
| 8 |
+
tags:
|
| 9 |
+
- vision-language-model
|
| 10 |
+
- structured-extraction
|
| 11 |
+
- table-extraction
|
| 12 |
+
- financial-tables
|
| 13 |
+
- qwen3-vl
|
| 14 |
+
- vlm
|
| 15 |
+
- dora
|
| 16 |
+
library_name: transformers
|
| 17 |
+
---
|
| 18 |
+
|
| 19 |
+
# structured-extractor-qwen3vl-4b-exp93
|
| 20 |
+
|
| 21 |
+
A fine-tuned **Qwen3-VL-4B-Instruct** that extracts structured rows
|
| 22 |
+
`(name, value, date, unit)` from images of financial-statement tables
|
| 23 |
+
(Russian + English). This is the **exp93** checkpoint — the project's
|
| 24 |
+
*reproducible* gold recipe across multiple seeds.
|
| 25 |
+
|
| 26 |
+
## Benchmarks
|
| 27 |
+
|
| 28 |
+
Evaluated on a held-out test split of real financial-statement table crops
|
| 29 |
+
(no synthetic data in training), with the `quality` preset
|
| 30 |
+
(`num_beams=4, repetition_penalty=1.1, min_new_tokens=200`):
|
| 31 |
+
|
| 32 |
+
| Metric | STRICT | LENIENT* |
|
| 33 |
+
|---|---:|---:|
|
| 34 |
+
| **tuple_f1** | **0.5316** | **0.5686** |
|
| 35 |
+
|
| 36 |
+
\* "LENIENT" normalizes unit synonyms (`million` ↔ `millions`, `млн руб.` ↔ `млн руб`) and accepts a year match when reference and prediction differ only by date precision. See `score_lenient.py` in this repo.
|
| 37 |
+
|
| 38 |
+
For comparison, on the same eval:
|
| 39 |
+
|
| 40 |
+
| Variant | STRICT | LENIENT |
|
| 41 |
+
|---|---:|---:|
|
| 42 |
+
| exp93 b4+rp1.1 (this repo) | 0.5316 | 0.5686 |
|
| 43 |
+
| exp93 greedy | ~0.50 | ~0.55 |
|
| 44 |
+
| exp85 greedy (DoRA+MLP without `wd=0.05`) | 0.5154 | — |
|
| 45 |
+
|
| 46 |
+
A best-of-3-seeds checkpoint (exp106) reaches LENIENT 0.5923, but it's a
|
| 47 |
+
*lucky* training trajectory — not reproducible from this recipe alone, so
|
| 48 |
+
this repo ships the recipe-reproducible exp93 instead.
|
| 49 |
+
|
| 50 |
+
> ⚠️ Earlier versions of this project reported t_f1 ~0.82 — those numbers
|
| 51 |
+
> were inflated by a target-leakage bug in the eval pipeline (the answer was
|
| 52 |
+
> in the model's input). The numbers above are real zero-shot, measured with
|
| 53 |
+
> a leak-free eval (`PageDataset(..., eval_mode=True)`).
|
| 54 |
+
|
| 55 |
+
## Recipe
|
| 56 |
+
|
| 57 |
+
- **Base**: `Qwen/Qwen3-VL-4B-Instruct` (Apache-2.0)
|
| 58 |
+
- **Adapter**: DoRA-style LoRA, `r=16`, `alpha=32`, `dropout=0.05`
|
| 59 |
+
- **Targets**: `q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj` (attention **+ MLP**)
|
| 60 |
+
- **Training data**: real financial tables only (~1.5k train, no synthetic augmentation)
|
| 61 |
+
- **Input**: pre-cropped table image + markdown OCR of that table + date-column hint
|
| 62 |
+
- **Schedule**: 2 epochs, AdamW, `lr=1e-4`, `weight_decay=0.05`, `warmup_ratio=0.05`
|
| 63 |
+
- **Saved as a full merged model** (8.3 GB safetensors), not a PEFT adapter.
|
| 64 |
+
|
| 65 |
+
## Quick start
|
| 66 |
+
|
| 67 |
+
```bash
|
| 68 |
+
pip install -r requirements.txt
|
| 69 |
+
```
|
| 70 |
+
|
| 71 |
+
```python
|
| 72 |
+
from inference import StructuredExtractor
|
| 73 |
+
|
| 74 |
+
extractor = StructuredExtractor.from_pretrained(
|
| 75 |
+
"Glazkov/structured-extractor-qwen3vl-4b-exp93"
|
| 76 |
+
)
|
| 77 |
+
|
| 78 |
+
result = extractor.extract(
|
| 79 |
+
"table_crop.png",
|
| 80 |
+
markdown=table_md_text,
|
| 81 |
+
date_columns=["2024", "2023"],
|
| 82 |
+
preset="quality",
|
| 83 |
+
)
|
| 84 |
+
|
| 85 |
+
for row in result["parameters"]:
|
| 86 |
+
print(row)
|
| 87 |
+
# {"parameter_name": "Interest income", "parameter_value": "533",
|
| 88 |
+
# "parameter_date": "2024", "parameter_unit": "millions"}
|
| 89 |
+
# ...
|
| 90 |
+
```
|
| 91 |
+
|
| 92 |
+
## Required inputs
|
| 93 |
+
|
| 94 |
+
The model was trained with three input streams. All three matter:
|
| 95 |
+
|
| 96 |
+
| Input | Status | Notes |
|
| 97 |
+
|---|---|---|
|
| 98 |
+
| **Table image** | Required | Pre-cropped to a single table region; long-side resized to 1344px (handled internally) |
|
| 99 |
+
| **Markdown OCR** of that table | Strongly recommended | The per-sample disambiguator on multi-table pages. Without it the model picks an arbitrary table and tuple-F1 drops to near zero. |
|
| 100 |
+
| **`date_columns`** hint | Optional | List of date-column headers; helps when markdown is noisy |
|
| 101 |
+
|
| 102 |
+
**The table image must be cropped to the target table**, not a full page.
|
| 103 |
+
The training data uses single-table crops; full-page images at inference are untested and likely degrade quality.
|
| 104 |
+
|
| 105 |
+
## Best-quality pipeline
|
| 106 |
+
|
| 107 |
+
```python
|
| 108 |
+
result = extractor.extract(
|
| 109 |
+
"table_crop.png",
|
| 110 |
+
markdown=table_md_text,
|
| 111 |
+
date_columns=["2024", "2023"],
|
| 112 |
+
preset="quality",
|
| 113 |
+
)
|
| 114 |
+
```
|
| 115 |
+
|
| 116 |
+
`preset="quality"` is `num_beams=4, length_penalty=1.0, min_new_tokens=200,
|
| 117 |
+
repetition_penalty=1.1, max_new_tokens=4096`. This is the configuration that
|
| 118 |
+
yields STRICT 0.5316 / LENIENT 0.5686.
|
| 119 |
+
|
| 120 |
+
## Most-optimal pipeline (greedy)
|
| 121 |
+
|
| 122 |
+
```python
|
| 123 |
+
result = extractor.extract(
|
| 124 |
+
"table_crop.png",
|
| 125 |
+
markdown=table_md_text,
|
| 126 |
+
date_columns=["2024", "2023"],
|
| 127 |
+
preset="fast",
|
| 128 |
+
)
|
| 129 |
+
```
|
| 130 |
+
|
| 131 |
+
Greedy decoding (`num_beams=1`). About **3-4× faster** than `quality` with a
|
| 132 |
+
~0.03-0.05 t_f1 drop. Use this when latency or throughput matters more than
|
| 133 |
+
the last point of F1.
|
| 134 |
+
|
| 135 |
+
## Batch inference
|
| 136 |
+
|
| 137 |
+
```python
|
| 138 |
+
from pathlib import Path
|
| 139 |
+
from inference import StructuredExtractor
|
| 140 |
+
|
| 141 |
+
extractor = StructuredExtractor.from_pretrained(
|
| 142 |
+
"Glazkov/structured-extractor-qwen3vl-4b-exp93"
|
| 143 |
+
)
|
| 144 |
+
|
| 145 |
+
paths = sorted(Path("tables/").glob("*.png"))
|
| 146 |
+
markdowns = [Path(p.with_suffix(".md")).read_text() for p in paths]
|
| 147 |
+
results = extractor.extract_batch(
|
| 148 |
+
paths,
|
| 149 |
+
markdown_batch=markdowns,
|
| 150 |
+
preset="fast",
|
| 151 |
+
batch_size=1, # beam search is memory-hungry; keep at 1
|
| 152 |
+
)
|
| 153 |
+
```
|
| 154 |
+
|
| 155 |
+
See `examples/batch.py` for a CLI version. `batch_size>1` is unsupported in
|
| 156 |
+
this wrapper because beam-search batching requires the training-time
|
| 157 |
+
collator (left-padding + cat of vision tensors) which is out of scope for an
|
| 158 |
+
inference module.
|
| 159 |
+
|
| 160 |
+
## Lenient scoring helper
|
| 161 |
+
|
| 162 |
+
`score_lenient.py` re-scores a JSONL of `(image, parameters)` predictions
|
| 163 |
+
against a reference annotations JSONL using unit aliases and date-year
|
| 164 |
+
normalization. The +0.037 LENIENT lift in the benchmark table comes from
|
| 165 |
+
this scorer; the model output itself is identical.
|
| 166 |
+
|
| 167 |
+
```bash
|
| 168 |
+
python score_lenient.py preds.jsonl annotations_test.jsonl
|
| 169 |
+
```
|
| 170 |
+
|
| 171 |
+
## Output format
|
| 172 |
+
|
| 173 |
+
The model emits one parameter per line in pipe-separated `sep_labels` format:
|
| 174 |
+
|
| 175 |
+
```
|
| 176 |
+
<|sep_meta|>
|
| 177 |
+
name: Interest income|value: 533|date: 2024|unit: millions
|
| 178 |
+
name: Foreign-currency transaction loss|value: 89|date: 2023|unit: millions
|
| 179 |
+
```
|
| 180 |
+
|
| 181 |
+
`parser.py` (in this repo) converts that to `{"parameters": [{...}, ...]}`.
|
| 182 |
+
**The parser strips stray `<|...|>` control-token artifacts before splitting** —
|
| 183 |
+
the model occasionally emits one mid-row, and without this strip a leading
|
| 184 |
+
`<` contaminates the previous field. This fix is worth +0.024-0.042 t_f1 on
|
| 185 |
+
its own.
|
| 186 |
+
|
| 187 |
+
## Loading details
|
| 188 |
+
|
| 189 |
+
`StructuredExtractor.from_pretrained` does three things you'd otherwise need
|
| 190 |
+
to wire up yourself:
|
| 191 |
+
|
| 192 |
+
1. Loads the processor (image processor + chat template) — first tries the
|
| 193 |
+
uploaded checkpoint, falls back to `Qwen/Qwen3-VL-4B-Instruct` if the
|
| 194 |
+
preprocessor configs aren't present.
|
| 195 |
+
2. Swaps in the fine-tuned tokenizer (which has the 4 added special tokens:
|
| 196 |
+
`<|sep_meta|>`, `<|sep_columns|>`, `<|sep_rows|>`, `<|sep_end|>`).
|
| 197 |
+
3. Force-injects `<|sep_meta|>\n` as the assistant-turn prefix before
|
| 198 |
+
`generate()`. This token is masked out of training labels — the model
|
| 199 |
+
never learned to emit it, so we have to prime it.
|
| 200 |
+
|
| 201 |
+
## Hardware
|
| 202 |
+
|
| 203 |
+
| Preset | Min VRAM (single image) |
|
| 204 |
+
|---|---:|
|
| 205 |
+
| fast (greedy) | ~12 GB |
|
| 206 |
+
| quality (beam=4) | ~24 GB |
|
| 207 |
+
|
| 208 |
+
bf16 on CUDA capability ≥ 8.0, float16 elsewhere. CPU works but is
|
| 209 |
+
unusably slow for a 4B VLM with beam search.
|
| 210 |
+
|
| 211 |
+
## Limitations
|
| 212 |
+
|
| 213 |
+
- Trained on financial-statement tables (RU/EN). Behavior on other domains
|
| 214 |
+
is unmeasured.
|
| 215 |
+
- **Bimodal errors**: ~42% of test samples solve well (t_f1 ≥ 0.7), ~34%
|
| 216 |
+
fail completely (t_f1 < 0.1). Worst failures cluster in specific source
|
| 217 |
+
documents (multi-table pages where the markdown disambiguator alone isn't
|
| 218 |
+
enough).
|
| 219 |
+
- Single-seed numbers vary a lot on this small dataset (~0.20 range across
|
| 220 |
+
4 seeds). exp93 is reproducible, but don't expect another retrain to
|
| 221 |
+
land at exactly the same F1.
|
| 222 |
+
|
| 223 |
+
## License
|
| 224 |
+
|
| 225 |
+
Apache-2.0, matching the base model.
|
| 226 |
+
|
| 227 |
+
## Citation / acknowledgements
|
| 228 |
+
|
| 229 |
+
- Base model: [`Qwen/Qwen3-VL-4B-Instruct`](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct) (Apache-2.0)
|
| 230 |
+
- Training framework: `structured-extractor-train` (DoRA r=16 + MLP, 2 epochs, real-only)
|