[NeurIPS 2023 WS] An Evaluator LM that is open-source, offers reproducible evaluation, and inexpensive to use. Specifically designed for fine-grained evaluation on a customized score rubric, Prometheus is a good alternative for human evaluation and GPT-4 evaluation.
active 2023-10-15 → 2024-04-29 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 198 days from 2023-10-15 to 2024-04-29. Pushes: 16 total, peak 14 in a day. Pull requests: 3 total, peak 2 in a day. Issues: 26 total, peak 7 in a day. Comments: 21 total, peak 5 in a day. Stars: 231 total, peak 29 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| SeungoneKim | 35 | 8 | 3 | 14 |
| jshin49 | 8 | 8 | 0 | 0 |
| HuihuiChyan | 4 | 0 | 0 | 2 |
| shaoyijia | 3 | 0 | 0 | 2 |
| mqo00 | 2 | 0 | 0 | 1 |
| ChiaraOleary | 2 | 0 | 0 | 0 |
| je1lee | 1 | 0 | 0 | 0 |
| agokrani | 1 | 0 | 0 | 1 |
| se-ok | 1 | 0 | 0 | 0 |
| zhao1402072392 | 1 | 0 | 0 | 0 |
| Haoxiang-Wang | 1 | 0 | 0 | 0 |
| ogencoglu | 1 | 0 | 0 | 1 |
| nnethercott | 1 | 0 | 0 | 0 |
| deshwalmahesh | 1 | 0 | 0 | 0 |
| maurovitaleBH | 1 | 0 | 0 | 0 |
| WoutDeRijck | 1 | 0 | 0 | 0 |
| sungkim11 | 1 | 0 | 0 | 0 |
| gmftbyGMFTBY | 1 | 0 | 0 | 0 |
Recent activity
Latest issues, pull requests and releases
- Issue#16se-ok2024-04-04 11:16ood_test missing some gpt4 feedback
- Issue comment#15ogencoglu2024-02-23 06:32Version Issue for BetterTransformer. Please provide exact package dependencies and Python, Torch version you used
- Issue#14SeungoneKim2024-02-16 15:20Prometheus using no reference materials
- Issue comment#14SeungoneKim2024-02-13 20:40Prometheus using no reference materials
- Issue#15deshwalmahesh2024-02-12 18:19Version Issue for BetterTransformer. Please provide exact package dependencies and Python, Torch version you used
- Issue#14maurovitaleBH2024-02-12 17:13Prometheus using no reference materials
- Issue#13SeungoneKim2024-01-16 04:04Demo of Prometheus
- Issue comment#12SeungoneKim2024-01-16 04:02Unable to generate evaluation
- Issue#12HuihuiChyan2024-01-16 03:48Unable to generate evaluation
- Issue comment#12HuihuiChyan2024-01-16 03:48Unable to generate evaluation
- Issue comment#12HuihuiChyan2024-01-16 03:04Unable to generate evaluation
- Issue comment#13SeungoneKim2024-01-15 04:47Demo of Prometheus
- Issue comment#12SeungoneKim2024-01-15 04:45Unable to generate evaluation
- Issue#13zhao14020723922024-01-12 09:25Demo of Prometheus
- Issue#12HuihuiChyan2024-01-11 04:29Unable to generate evaluation
- Issue comment#7mqo002023-12-31 03:21A functional command example for model training (also on a single GPU)?
- Issue#11SeungoneKim2023-12-26 04:17Question About Feedback Bench
- Issue comment#11SeungoneKim2023-12-26 04:17Question About Feedback Bench
- Issue#10SeungoneKim2023-12-26 04:15Grad clipping for fp16
- Issue comment#10SeungoneKim2023-12-26 04:15Grad clipping for fp16
- Issue#9SeungoneKim2023-12-26 04:12score rubric label for feedback_collection_test.json
- Issue comment#9SeungoneKim2023-12-26 04:11score rubric label for feedback_collection_test.json
- Issue#8SeungoneKim2023-12-26 04:08Evaluation Code for AlpacaFarm, FLASK, MT-Bench
- Issue comment#8SeungoneKim2023-12-26 04:07Evaluation Code for AlpacaFarm, FLASK, MT-Bench
- Issue#7SeungoneKim2023-12-26 04:05A functional command example for model training (also on a single GPU)?
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 231 stars here means stars gained during the window, not the repo's star count.