ML / OCR Engineer Paid Internship
We are looking for an ML/OCR engineering intern to help digitize printed Amharic and Ge’ez book pages into structured data.
You will work on OCR, computer vision, and human-in-the-loop annotation systems for low-resource scripts. This is a hands-on role for someone comfortable experimenting with models, evaluating results, and improving a real pipeline over time.
What you’ll do:
* Build and fine-tune OCR models for printed Amharic and Ge’ez textImprove text detection, recognition, and page layout analysis
* Train object detection and classification models for document-processing tasks
* Help design a human-in-the-loop annotation workflow, including active learning and error correction loops
* Evaluate model performance, benchmark experiments, and ship measurable improvements each cycle
* Use vision-language models or LLM APIs where helpful for weak supervision, annotation drafting, or quality checks.
Requirement
* Strong Python skills
* Experience training machine learning or computer vision models.
* Familiarity with PyTorch or TensorFlow.
* Interest in OCR, document understanding, or low-resource language processing.
* Ability to run experiments, analyze errors, and communicate results clearly.
Nice to have
* Prior OCR experience, especially text detection, recognition, or layout analysis
* Experience with Tesseract fine-tuning, TrOCR, PaddleOCR, EasyOCR, or custom CTC/transformer OCR pipelines
* Experience with object detection models such as YOLO, DETR, or similar architectures
* Experience working with low-resource scripts or languages
* Familiarity with active learning or human-in-the-loop annotation tools
* Familiarity with vision LLM APIs for drafting, weak supervision, or validation
* Audio ML experience, such as segmentation, forced alignment, or Whisper
Send your resume at: [email protected]
14
13June 11, 2026 3.2K 63