Skip to content

[NeurIPS 2023 WS] An Evaluator LM that is open-source, offers reproducible evaluation, and inexpensive to use. Specifically designed for fine-grained evaluation on a customized score rubric, Prometheus is a good alternative for human evaluation and GPT-4 evaluation.

active 2023-10-152024-04-29 (UTC)

Complete coverage26,587 / 26,587 hourly files (100%) · 2 absent upstream2023-08-152026-08-26 (UTC)
Events
319
Pushes
16
Pull requests
3
Issues
26
Stars
231
Forks
21

Activity over time

Daily event counts in the loaded window

Line chart, 198 days from 2023-10-15 to 2024-04-29. Pushes: 16 total, peak 14 in a day. Pull requests: 3 total, peak 2 in a day. Issues: 26 total, peak 7 in a day. Comments: 21 total, peak 5 in a day. Stars: 231 total, peak 29 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

Recent activity

Latest issues, pull requests and releases

  • Issue#16se-ok2024-04-04 11:16
    ood_test missing some gpt4 feedback
  • Issue comment#15ogencoglu2024-02-23 06:32
    Version Issue for BetterTransformer. Please provide exact package dependencies and Python, Torch version you used
  • Issue#14SeungoneKim2024-02-16 15:20
    Prometheus using no reference materials
  • Issue comment#14SeungoneKim2024-02-13 20:40
    Prometheus using no reference materials
  • Issue#15deshwalmahesh2024-02-12 18:19
    Version Issue for BetterTransformer. Please provide exact package dependencies and Python, Torch version you used
  • Issue#14maurovitaleBH2024-02-12 17:13
    Prometheus using no reference materials
  • Issue#13SeungoneKim2024-01-16 04:04
    Demo of Prometheus
  • Issue comment#12SeungoneKim2024-01-16 04:02
    Unable to generate evaluation
  • Issue#12HuihuiChyan2024-01-16 03:48
    Unable to generate evaluation
  • Issue comment#12HuihuiChyan2024-01-16 03:48
    Unable to generate evaluation
  • Issue comment#12HuihuiChyan2024-01-16 03:04
    Unable to generate evaluation
  • Issue comment#13SeungoneKim2024-01-15 04:47
    Demo of Prometheus
  • Issue comment#12SeungoneKim2024-01-15 04:45
    Unable to generate evaluation
  • Issue#13zhao14020723922024-01-12 09:25
    Demo of Prometheus
  • Issue#12HuihuiChyan2024-01-11 04:29
    Unable to generate evaluation
  • Issue comment#7mqo002023-12-31 03:21
    A functional command example for model training (also on a single GPU)?
  • Issue#11SeungoneKim2023-12-26 04:17
    Question About Feedback Bench
  • Issue comment#11SeungoneKim2023-12-26 04:17
    Question About Feedback Bench
  • Issue#10SeungoneKim2023-12-26 04:15
    Grad clipping for fp16
  • Issue comment#10SeungoneKim2023-12-26 04:15
    Grad clipping for fp16
  • Issue#9SeungoneKim2023-12-26 04:12
    score rubric label for feedback_collection_test.json
  • Issue comment#9SeungoneKim2023-12-26 04:11
    score rubric label for feedback_collection_test.json
  • Issue#8SeungoneKim2023-12-26 04:08
    Evaluation Code for AlpacaFarm, FLASK, MT-Bench
  • Issue comment#8SeungoneKim2023-12-26 04:07
    Evaluation Code for AlpacaFarm, FLASK, MT-Bench
  • Issue#7SeungoneKim2023-12-26 04:05
    A functional command example for model training (also on a single GPU)?

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 231 stars here means stars gained during the window, not the repo's star count.