Skip to content

๐Ÿค— Evaluate: A library for easily evaluating machine learning models and datasets.

Python ยท active 2024-01-12 โ†’ 2026-07-06 (UTC)

Partial coverage17,627 / 22,635 hourly files (78%) ยท 2 absent upstream ยท 5,008 failed, retryable2024-01-12 โ†’ 2026-08-12 (UTC)โ€” sampled evenly across the window, so rankings and trends hold; absolute counts scale up.
Events
753
Pushes
27
Pull requests
48
Issues
56
Stars
396
Forks
70

Activity over time

Daily event counts in the loaded window

Line chart, 907 days from 2024-01-12 to 2026-07-06. Pushes: 27 total, peak 7 in a day. Pull requests: 48 total, peak 4 in a day. Issues: 56 total, peak 3 in a day. Comments: 122 total, peak 7 in a day. Stars: 396 total, peak 5 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 โ€” โˆ’95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments โ€” stars and forks excluded, so this is contribution rather than popularity

Recent activity

Latest issues, pull requests and releases

  • Pull request#775zanvari2026-07-05 00:24
  • Pull request#763connerlambden2026-06-05 04:33
  • Issue comment#165cakiki2026-04-27 18:42
    Add common metrics for information retrieval
  • Issue comment#740plutonium-2392026-04-14 09:41
    auto-update metrics spaces to latest gradio versions with action; fixes #713
  • Pull request#744paulinebm2026-04-02 09:31
  • Pull request#743SumitVermakgp2026-04-01 23:35
  • Issue#742SumitVermakgp2026-04-01 23:26
    New community metric: RAIL Score โ€” responsible AI evaluation across 8 dimensions
  • Pull request#738joshuaswanson2026-03-12 12:25
  • Pull request#734aakash-agarwal-0022026-02-16 17:28
  • Pull request#733michaelellis0032026-02-13 14:32
  • Pull request#732Ashutosh0x2026-02-08 11:04
  • Issue comment#551RylanSchaeffer2026-02-01 20:55
    Add geometric mean of per-token-Perplexities
  • Pull request#730Vangmay2026-01-28 01:01
  • Pull request#728dyra-122026-01-26 16:56
  • Pull request#727burtenshaw2026-01-20 17:12
  • Issue comment#671521472026-01-05 00:55
    Suppress the output of nltk download logs when loading metric via evaluate
  • Pull request#726521472026-01-05 00:54
  • Issue#724Ayesha-Imr2025-12-27 19:43
    Proposal: Make SQuAD QA F1 logic reusable as a generic QA metric in `evaluate`
  • Issue#723viviansmlie2025-12-18 06:49
    FileNotFoundError: Couldn't find a module script at /workspace/weiw13@xiaopeng.com/VLM/evaluate/meteor/meteor.py. Module 'meteor' doesn't exist on the Hugging Face Hub either.
  • Pull request#721ajeetkartikay2025-12-10 20:22
  • Pull request#720ajeetkartikay2025-12-10 19:18
  • Issue comment#717lhoestq2025-11-14 14:39
    Fix dependency hints on ImportError
  • Pull request#718Yacklin2025-11-14 01:53
  • Issue comment#716Yacklin2025-11-14 01:35
    Remove the comments that pollute library_import_path in _download_additional_modules and correct import_library_name for "absl"
  • Issue comment#716Yacklin2025-11-13 17:12
    Remove the comments that pollute library_import_path in _download_additional_modules and correct import_library_name for "absl"

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals โ€” 396 stars here means stars gained during the window, not the repo's star count.