Skip to content

📝 针对文档类图像做内容提取,将文档类图像一比一输出到Word或者Txt中,便于进一步使用或处理。后续计划支持输入PDF/图像,输出对应json格式、Txt格式、Word格式和Markdown格式。

Python · active 2024-10-112026-06-28 (UTC)

Complete coverage26,489 / 26,489 hourly files (100%) · 2 absent upstream2023-08-152026-08-22 (UTC)
Events
301
Pushes
70
Pull requests
2
Issues
30
Stars
113
Forks
10

Activity over time

Daily event counts in the loaded window

Line chart, 626 days from 2024-10-11 to 2026-06-28. Pushes: 70 total, peak 5 in a day. Pull requests: 2 total, peak 2 in a day. Issues: 30 total, peak 3 in a day. Comments: 69 total, peak 9 in a day. Stars: 113 total, peak 3 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
hzkitty9562030
nissansz320022
SWHL13624
Fusyong7003
hua35054002
azwhale3002
CY2022272001
lolocoo2002
znsoftm2200
lmw03201001
ningpp1000
DeoLeung1001
beetter1000
simjak1000
meimosor1000
ZenCoderWu1000
liang-xian1001
zhcn0000001000
asfoorial1000
heyuqi19701000

Recent activity

Latest issues, pull requests and releases

  • Issue comment#48DeoLeung2026-06-08 00:54
    放松 pdfminer.six 和 pydantic 版本限制
  • Releasehzkitty2026-05-24 14:36
    v0.9.5
  • Issue comment#47hzkitty2026-05-10 16:00
    不支持rapidocr >= 3.4.3,希望尽快适配
  • Issue comment#43nissansz2026-04-12 22:14
    版面分析模型,想只获取段落文字,包括表格里单元格的段落, 用哪个模型好?
  • Issue comment#43hzkitty2026-04-12 18:10
    版面分析模型,想只获取段落文字,包括表格里单元格的段落, 用哪个模型好?
  • Issue comment#42hzkitty2026-04-12 18:08
    不知道是不是调用方式不对,测试结果有点奇怪
  • Issue comment#9nissansz2026-03-29 22:13
    如何单独调用layout 版面分析,返回结果?
  • Issue comment#9nissansz2026-03-28 23:53
    如何单独调用layout 版面分析,返回结果?
  • Issue comment#32Fusyong2026-03-17 06:31
    功能请求:在原PDF基础上生成可搜索的双层PDF
  • Issue#32Fusyong2026-03-17 03:38
    功能请求:在原PDF基础上生成可搜索的双层PDF
  • Issue comment#32Fusyong2026-03-17 03:38
    功能请求:在原PDF基础上生成可搜索的双层PDF
  • Issue comment#40nissansz2026-03-12 03:56
    内置 PaddleOCRVL 系列模型的集成 paddleocr_vl.py (ocr、formula、table), 这个也是onnx模型?怎么调用?
  • Issue#40nissansz2026-03-11 23:22
    内置 PaddleOCRVL 系列模型的集成 paddleocr_vl.py (ocr、formula、table), 这个也是onnx模型?怎么调用?
  • Issue#39zhcn0000002026-03-07 09:13
    能否以 markdown 原生的形式输出表格,代码块等而不是使用 html 标签
  • Issue#37asfoorial2026-02-15 09:30
    TableMatch.get_pred_html is not adding spaces between words in a table
  • Issue comment#36lmw03202026-02-10 09:51
    使用PPlayoutV2的自定义模型路径时报错
  • Releasehzkitty2026-02-08 16:11
    v0.7.0
  • Issue#32Fusyong2026-01-05 03:22
    功能请求:在原PDF基础上生成可搜索的双层PDF
  • Issue comment#32Fusyong2026-01-05 03:22
    功能请求:在原PDF基础上生成可搜索的双层PDF
  • Issue comment#34nissansz2026-01-03 14:51
    "dots.ocr-1.8B-Q8_0.gguf" 这种模型文件怎么用python调用?
  • Issue comment#32hzkitty2026-01-02 18:59
    功能请求:在原PDF基础上生成可搜索的双层PDF
  • Issue comment#32hzkitty2026-01-02 18:43
    功能请求:在原PDF基础上生成可搜索的双层PDF
  • Issue comment#33hzkitty2026-01-02 16:51
    PP-DocLayoutV2模型批量推理报错
  • Issue#33ningpp2026-01-02 06:52
    PP-DocLayoutV2模型批量推理报错
  • Issue#32Fusyong2026-01-02 03:53
    功能请求:在原PDF基础上生成可搜索的双层PDF

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 113 stars here means stars gained during the window, not the repo's star count.