Request for Qari-OCR training data download links and license clarification

#1
by cloudaocr - opened

Hello Qari-OCR team,

I am developing an Arabic OCR system for printed, historical, and low-quality scanned documents.

Could you please provide the official download links for the training, validation, and test data used for Qari-OCR 0.4.0, including:

  • page images or PDF files;
  • corresponding ground-truth/reference text;
  • image-to-text mapping files;
  • annotations such as JSONL, XML, Markdown, ALTO, PAGE-XML, or metadata;
  • the official train, validation, and test splits.

I would also appreciate written clarification on whether I may:

  1. use the data to train or fine-tune another OCR or vision-language model;
  2. create internal augmented or degraded copies;
  3. publish benchmark results and training/integration code;
  4. publish LoRA adapters or trained/merged model weights;
  5. use the resulting model in a future commercial product or paid API;
  6. retain the original data privately without redistributing it.

Please also clarify whether any source datasets used for Qari-OCR 0.4.0 have separate restrictions and what attribution is required.

I previously contacted the team by email and am posting here in case this is the preferred communication channel.

Best regards,
Ahmad Sayed

Network for Advancing Modern ArabicNLP & AI org

Hi Ahmad,
Thanks for your interest in Qari-OCR.

The Qari-OCR model is released under the Apache 2.0 license, which allows both research and commercial use, including fine-tuning, creating derivative models, publishing adapters or model weights, and integrating the model into commercial applications, subject to the terms of the Apache 2.0 license.

Regarding the datasets you mentioned (training, validation, and test data), could you clarify what you are referring to? Your question appears to be about the datasets rather than the Qari-OCR model itself.

Best,
Qari-OCR Team

Hi,

Thank you for the clarification.

By “datasets,” I mean the data used to train, validate, and evaluate Qari-OCR itself, including:

  • the page images;
  • the corresponding ground-truth/reference text;
  • image-to-text mapping files;
  • train, validation, and test splits;
  • dataset format and documentation;
  • the license that applies to each dataset.

Could you please share the official download links or repository names for these datasets?

I would also like to confirm whether I may use these datasets to:

  1. train or fine-tune another OCR or vision-language model;
  2. create internal augmented or degraded copies;
  3. publish benchmark results and training code;
  4. publish LoRA adapters or trained/merged model weights;
  5. use the resulting model in a commercial product or paid API;
  6. retain the original datasets privately without redistributing them.

Please also clarify whether any source images, books, or third-party assets have separate restrictions.

Best regards,
Ahmad Sayed

Network for Advancing Modern ArabicNLP & AI org

we do not provide the full dataset .. but you may find some in our repo

thank you

Omartificial-Intelligence-Space changed discussion status to closed

Sign up or log in to comment