Initial commit: Invoice field extraction system using YOLO + OCR
Features: - Auto-labeling pipeline: CSV values -> PDF search -> YOLO annotations - Flexible date matching: year-month match, nearby date tolerance - PDF text extraction with PyMuPDF - OCR support for scanned documents (PaddleOCR) - YOLO training and inference pipeline - 7 field types: InvoiceNumber, InvoiceDate, InvoiceDueDate, OCR, Bankgiro, Plusgiro, Amount Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
This commit is contained in:
3
src/ocr/__init__.py
Normal file
3
src/ocr/__init__.py
Normal file
@@ -0,0 +1,3 @@
|
||||
from .paddle_ocr import OCREngine, extract_ocr_tokens
|
||||
|
||||
__all__ = ['OCREngine', 'extract_ocr_tokens']
|
||||
Reference in New Issue
Block a user