Skip to content

OpenPecha/Create_OCR_benchmark_data

View on GitHub ↗Related repositories →

active 2024-03-012024-05-14 (UTC)

Partial coverage17,185 / 21,947 hourly files (78%) · 2 absent upstream · 4,761 failed, retryable2024-02-092026-08-12 (UTC)— sampled evenly across the window, so rankings and trends hold; absolute counts scale up.
Events
17
Pushes
1
Pull requests
0
Issues
13
Stars
0
Forks
0

Activity over time

Daily event counts in the loaded window

Line chart, 75 days from 2024-03-01 to 2024-05-14. Pushes: 1 total, peak 1 in a day. Pull requests: 0 total, peak 0 in a day. Issues: 13 total, peak 6 in a day. Comments: 1 total, peak 1 in a day. Stars: 0 total, peak 0 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
ta4tsering13100
tenzin32001

Recent activity

Latest issues, pull requests and releases

  • Issue comment#7tenzin32024-05-14 05:02
    OCR0023: Create Google Books benchmark set.
  • Issue#1ta4tsering2024-05-13 05:43
    Download all the annotation jsonl from batch 1 - 18 and norbuletaka batch 1-150.
  • Issue#2ta4tsering2024-05-13 05:43
    Create a csv for each batch.
  • Issue#3ta4tsering2024-05-13 05:43
    Randomly select 10% of images from each batch.
  • Issue#4ta4tsering2024-05-13 05:43
    create csv for each transcripts of the training data and remove the above selected images.
  • Issue#5ta4tsering2024-05-13 05:43
    Delete the benchmark data images from the training data images.
  • Issue#6ta4tsering2024-05-13 05:43
    Create a new folder with all the OCR training data in zip
  • Issue#7tenzin32024-05-10 04:40
    OCR0023: Create Google Books benchmark set.
  • Issue#6ta4tsering2024-03-01 07:32
    Create a new folder with all the OCR training data in zip
  • Issue#5ta4tsering2024-03-01 07:32
    Delete the benchmark data images from the training data images.
  • Issue#4ta4tsering2024-03-01 07:32
    create csv for each transcripts of the training data and remove the above selected images.
  • Issue#3ta4tsering2024-03-01 07:31
    Randomly select 10% of images from each batch.
  • Issue#2ta4tsering2024-03-01 07:31
    Create a csv for each batch.
  • Issue#1ta4tsering2024-03-01 07:31
    Download all the annotation jsonl from batch 1 - 18 and norbuletaka batch 1-150.

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 0 stars here means stars gained during the window, not the repo's star count.