tesseract-ocr/tessdata ? reverse-engineered prompt
Reverse engineered prompt
Build me a simple repository for Tesseract OCR language data files.
I want it to act like a clean package of trained models for Tesseract 4 and newer, with the fast integerized LSTM models and the older legacy models where they still make sense. It should include the language files for the common scripts and languages, plus the orientation and script detection model, and it should be ready to drop into a Tesseract install.
Please add a straightforward README that explains what these files are for, which Tesseract versions they work with, and the difference between the faster models and the more accurate ones. Also include a basic license file and any small helper configs that make it easy to use the trained data correctly.
If there are any current details that need checking against the latest Tesseract docs, look them up online and make sure the instructions are up to date.
Are you gonna build this?
make sure you review the code using coderabbit