Glazkov commited on
Commit
aa23949
·
verified ·
1 Parent(s): f371181

Add README.md

Browse files
Files changed (1) hide show
  1. README.md +230 -0
README.md ADDED
@@ -0,0 +1,230 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ - ru
6
+ base_model: Qwen/Qwen3-VL-4B-Instruct
7
+ pipeline_tag: image-text-to-text
8
+ tags:
9
+ - vision-language-model
10
+ - structured-extraction
11
+ - table-extraction
12
+ - financial-tables
13
+ - qwen3-vl
14
+ - vlm
15
+ - dora
16
+ library_name: transformers
17
+ ---
18
+
19
+ # structured-extractor-qwen3vl-4b-exp93
20
+
21
+ A fine-tuned **Qwen3-VL-4B-Instruct** that extracts structured rows
22
+ `(name, value, date, unit)` from images of financial-statement tables
23
+ (Russian + English). This is the **exp93** checkpoint — the project's
24
+ *reproducible* gold recipe across multiple seeds.
25
+
26
+ ## Benchmarks
27
+
28
+ Evaluated on a held-out test split of real financial-statement table crops
29
+ (no synthetic data in training), with the `quality` preset
30
+ (`num_beams=4, repetition_penalty=1.1, min_new_tokens=200`):
31
+
32
+ | Metric | STRICT | LENIENT* |
33
+ |---|---:|---:|
34
+ | **tuple_f1** | **0.5316** | **0.5686** |
35
+
36
+ \* "LENIENT" normalizes unit synonyms (`million` ↔ `millions`, `млн руб.` ↔ `млн руб`) and accepts a year match when reference and prediction differ only by date precision. See `score_lenient.py` in this repo.
37
+
38
+ For comparison, on the same eval:
39
+
40
+ | Variant | STRICT | LENIENT |
41
+ |---|---:|---:|
42
+ | exp93 b4+rp1.1 (this repo) | 0.5316 | 0.5686 |
43
+ | exp93 greedy | ~0.50 | ~0.55 |
44
+ | exp85 greedy (DoRA+MLP without `wd=0.05`) | 0.5154 | — |
45
+
46
+ A best-of-3-seeds checkpoint (exp106) reaches LENIENT 0.5923, but it's a
47
+ *lucky* training trajectory — not reproducible from this recipe alone, so
48
+ this repo ships the recipe-reproducible exp93 instead.
49
+
50
+ > ⚠️ Earlier versions of this project reported t_f1 ~0.82 — those numbers
51
+ > were inflated by a target-leakage bug in the eval pipeline (the answer was
52
+ > in the model's input). The numbers above are real zero-shot, measured with
53
+ > a leak-free eval (`PageDataset(..., eval_mode=True)`).
54
+
55
+ ## Recipe
56
+
57
+ - **Base**: `Qwen/Qwen3-VL-4B-Instruct` (Apache-2.0)
58
+ - **Adapter**: DoRA-style LoRA, `r=16`, `alpha=32`, `dropout=0.05`
59
+ - **Targets**: `q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj` (attention **+ MLP**)
60
+ - **Training data**: real financial tables only (~1.5k train, no synthetic augmentation)
61
+ - **Input**: pre-cropped table image + markdown OCR of that table + date-column hint
62
+ - **Schedule**: 2 epochs, AdamW, `lr=1e-4`, `weight_decay=0.05`, `warmup_ratio=0.05`
63
+ - **Saved as a full merged model** (8.3 GB safetensors), not a PEFT adapter.
64
+
65
+ ## Quick start
66
+
67
+ ```bash
68
+ pip install -r requirements.txt
69
+ ```
70
+
71
+ ```python
72
+ from inference import StructuredExtractor
73
+
74
+ extractor = StructuredExtractor.from_pretrained(
75
+ "Glazkov/structured-extractor-qwen3vl-4b-exp93"
76
+ )
77
+
78
+ result = extractor.extract(
79
+ "table_crop.png",
80
+ markdown=table_md_text,
81
+ date_columns=["2024", "2023"],
82
+ preset="quality",
83
+ )
84
+
85
+ for row in result["parameters"]:
86
+ print(row)
87
+ # {"parameter_name": "Interest income", "parameter_value": "533",
88
+ # "parameter_date": "2024", "parameter_unit": "millions"}
89
+ # ...
90
+ ```
91
+
92
+ ## Required inputs
93
+
94
+ The model was trained with three input streams. All three matter:
95
+
96
+ | Input | Status | Notes |
97
+ |---|---|---|
98
+ | **Table image** | Required | Pre-cropped to a single table region; long-side resized to 1344px (handled internally) |
99
+ | **Markdown OCR** of that table | Strongly recommended | The per-sample disambiguator on multi-table pages. Without it the model picks an arbitrary table and tuple-F1 drops to near zero. |
100
+ | **`date_columns`** hint | Optional | List of date-column headers; helps when markdown is noisy |
101
+
102
+ **The table image must be cropped to the target table**, not a full page.
103
+ The training data uses single-table crops; full-page images at inference are untested and likely degrade quality.
104
+
105
+ ## Best-quality pipeline
106
+
107
+ ```python
108
+ result = extractor.extract(
109
+ "table_crop.png",
110
+ markdown=table_md_text,
111
+ date_columns=["2024", "2023"],
112
+ preset="quality",
113
+ )
114
+ ```
115
+
116
+ `preset="quality"` is `num_beams=4, length_penalty=1.0, min_new_tokens=200,
117
+ repetition_penalty=1.1, max_new_tokens=4096`. This is the configuration that
118
+ yields STRICT 0.5316 / LENIENT 0.5686.
119
+
120
+ ## Most-optimal pipeline (greedy)
121
+
122
+ ```python
123
+ result = extractor.extract(
124
+ "table_crop.png",
125
+ markdown=table_md_text,
126
+ date_columns=["2024", "2023"],
127
+ preset="fast",
128
+ )
129
+ ```
130
+
131
+ Greedy decoding (`num_beams=1`). About **3-4× faster** than `quality` with a
132
+ ~0.03-0.05 t_f1 drop. Use this when latency or throughput matters more than
133
+ the last point of F1.
134
+
135
+ ## Batch inference
136
+
137
+ ```python
138
+ from pathlib import Path
139
+ from inference import StructuredExtractor
140
+
141
+ extractor = StructuredExtractor.from_pretrained(
142
+ "Glazkov/structured-extractor-qwen3vl-4b-exp93"
143
+ )
144
+
145
+ paths = sorted(Path("tables/").glob("*.png"))
146
+ markdowns = [Path(p.with_suffix(".md")).read_text() for p in paths]
147
+ results = extractor.extract_batch(
148
+ paths,
149
+ markdown_batch=markdowns,
150
+ preset="fast",
151
+ batch_size=1, # beam search is memory-hungry; keep at 1
152
+ )
153
+ ```
154
+
155
+ See `examples/batch.py` for a CLI version. `batch_size>1` is unsupported in
156
+ this wrapper because beam-search batching requires the training-time
157
+ collator (left-padding + cat of vision tensors) which is out of scope for an
158
+ inference module.
159
+
160
+ ## Lenient scoring helper
161
+
162
+ `score_lenient.py` re-scores a JSONL of `(image, parameters)` predictions
163
+ against a reference annotations JSONL using unit aliases and date-year
164
+ normalization. The +0.037 LENIENT lift in the benchmark table comes from
165
+ this scorer; the model output itself is identical.
166
+
167
+ ```bash
168
+ python score_lenient.py preds.jsonl annotations_test.jsonl
169
+ ```
170
+
171
+ ## Output format
172
+
173
+ The model emits one parameter per line in pipe-separated `sep_labels` format:
174
+
175
+ ```
176
+ <|sep_meta|>
177
+ name: Interest income|value: 533|date: 2024|unit: millions
178
+ name: Foreign-currency transaction loss|value: 89|date: 2023|unit: millions
179
+ ```
180
+
181
+ `parser.py` (in this repo) converts that to `{"parameters": [{...}, ...]}`.
182
+ **The parser strips stray `<|...|>` control-token artifacts before splitting** —
183
+ the model occasionally emits one mid-row, and without this strip a leading
184
+ `<` contaminates the previous field. This fix is worth +0.024-0.042 t_f1 on
185
+ its own.
186
+
187
+ ## Loading details
188
+
189
+ `StructuredExtractor.from_pretrained` does three things you'd otherwise need
190
+ to wire up yourself:
191
+
192
+ 1. Loads the processor (image processor + chat template) — first tries the
193
+ uploaded checkpoint, falls back to `Qwen/Qwen3-VL-4B-Instruct` if the
194
+ preprocessor configs aren't present.
195
+ 2. Swaps in the fine-tuned tokenizer (which has the 4 added special tokens:
196
+ `<|sep_meta|>`, `<|sep_columns|>`, `<|sep_rows|>`, `<|sep_end|>`).
197
+ 3. Force-injects `<|sep_meta|>\n` as the assistant-turn prefix before
198
+ `generate()`. This token is masked out of training labels — the model
199
+ never learned to emit it, so we have to prime it.
200
+
201
+ ## Hardware
202
+
203
+ | Preset | Min VRAM (single image) |
204
+ |---|---:|
205
+ | fast (greedy) | ~12 GB |
206
+ | quality (beam=4) | ~24 GB |
207
+
208
+ bf16 on CUDA capability ≥ 8.0, float16 elsewhere. CPU works but is
209
+ unusably slow for a 4B VLM with beam search.
210
+
211
+ ## Limitations
212
+
213
+ - Trained on financial-statement tables (RU/EN). Behavior on other domains
214
+ is unmeasured.
215
+ - **Bimodal errors**: ~42% of test samples solve well (t_f1 ≥ 0.7), ~34%
216
+ fail completely (t_f1 < 0.1). Worst failures cluster in specific source
217
+ documents (multi-table pages where the markdown disambiguator alone isn't
218
+ enough).
219
+ - Single-seed numbers vary a lot on this small dataset (~0.20 range across
220
+ 4 seeds). exp93 is reproducible, but don't expect another retrain to
221
+ land at exactly the same F1.
222
+
223
+ ## License
224
+
225
+ Apache-2.0, matching the base model.
226
+
227
+ ## Citation / acknowledgements
228
+
229
+ - Base model: [`Qwen/Qwen3-VL-4B-Instruct`](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct) (Apache-2.0)
230
+ - Training framework: `structured-extractor-train` (DoRA r=16 + MLP, 2 epochs, real-only)