因为10年前及以上出版的pdf书籍大多没有源文本,靠的是机器扫描成图片的形式做成的pdf。导致我想抓取文本做成本地知识库的时候,没有源数据。而传统的OCR识别正确率太低(问就是试过了依托答辩)所以做了个自动ai ocr的python脚本把数据识别抓取出来。 > > > > OCR(Optical Character Recognition,光学字符识别)是一种将图像中的文字转换为可编辑文本的技术。它广泛应用于文档数
active 2025-03-04 → 2025-03-04 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 1 days from 2025-03-04 to 2025-03-04. Pushes: 3 total, peak 3 in a day. Pull requests: 0 total, peak 0 in a day. Issues: 0 total, peak 0 in a day. Comments: 0 total, peak 0 in a day. Stars: 0 total, peak 0 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| buynonsense | 3 | 3 | 0 | 0 |
Recent activity
Latest issues, pull requests and releases
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 0 stars here means stars gained during the window, not the repo's star count.