Skip to content

buynonsense/PDFwithAIOCR

View on GitHub ↗Related repositories →

因为10年前及以上出版的pdf书籍大多没有源文本,靠的是机器扫描成图片的形式做成的pdf。导致我想抓取文本做成本地知识库的时候,没有源数据。而传统的OCR识别正确率太低(问就是试过了依托答辩)所以做了个自动ai ocr的python脚本把数据识别抓取出来。 > > > > OCR(Optical Character Recognition,光学字符识别)是一种将图像中的文字转换为可编辑文本的技术。它广泛应用于文档数

active 2025-03-04 → 2025-03-04 (UTC)

Complete coverage27,687 / 27,687 hourly files (100%) · 2 absent upstream2023-08-15 → 2026-10-11 (UTC)
Events
9
Pushes
3
Pull requests
0
Issues
0
Stars
0
Forks
0

Activity over time

Daily event counts in the loaded window

Line chart, 1 days from 2025-03-04 to 2025-03-04. Pushes: 3 total, peak 3 in a day. Pull requests: 0 total, peak 0 in a day. Issues: 0 total, peak 0 in a day. Comments: 0 total, peak 0 in a day. Stars: 0 total, peak 0 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
buynonsense3300

Recent activity

Latest issues, pull requests and releases

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 0 stars here means stars gained during the window, not the repo's star count.