Skip to content

Inspect: A framework for large language model evaluations

active 2024-05-082024-11-09 (UTC)

Partial coverage18,865 / 24,260 hourly files (78%) · 2 absent upstream · 5,393 failed, retryable2023-11-052026-08-12 (UTC)— sampled evenly across the window, so rankings and trends hold; absolute counts scale up.
Events
91
Pushes
44
Pull requests
11
Issues
0
Stars
0
Forks
2

Activity over time

Daily event counts in the loaded window

Line chart, 186 days from 2024-05-08 to 2024-11-09. Pushes: 44 total, peak 8 in a day. Pull requests: 11 total, peak 5 in a day. Issues: 0 total, peak 0 in a day. Comments: 14 total, peak 4 in a day. Stars: 0 total, peak 0 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
adil-a402974
XkunW9900
xeon277610
jjallaire-aisi7015
jjallaire6024
aisi-inspect1001

Recent activity

Latest issues, pull requests and releases

  • Pull request#7adil-a2024-08-09 01:31
  • Pull request#7adil-a2024-08-09 01:31
  • Pull request#6adil-a2024-08-09 01:30
  • Pull request#6adil-a2024-08-09 01:29
  • Pull request#1adil-a2024-08-09 00:30
  • Issue comment#1adil-a2024-08-08 20:19
    HumanEval Benchmark
  • Issue comment#1jjallaire2024-08-08 19:00
    HumanEval Benchmark
  • Issue comment#1jjallaire2024-08-07 22:34
    HumanEval Benchmark
  • Issue comment#1jjallaire-aisi2024-08-07 21:56
    HumanEval Benchmark
  • Issue comment#1jjallaire-aisi2024-08-07 21:54
    HumanEval Benchmark
  • Pull request#5jjallaire-aisi2024-08-07 21:51
  • Issue comment#1jjallaire2024-08-04 18:26
    HumanEval Benchmark
  • Issue comment#1adil-a2024-08-04 18:24
    HumanEval Benchmark
  • Issue comment#1aisi-inspect2024-08-04 18:09
    HumanEval Benchmark
  • Issue comment#1adil-a2024-08-04 17:57
    HumanEval Benchmark
  • Pull request#4adil-a2024-08-04 17:40
  • Pull request#4jjallaire2024-08-03 23:09
  • Issue comment#3jjallaire2024-07-30 21:28
    suggested improvements to humaneval
  • Pull request#3jjallaire2024-07-27 16:27
  • Issue comment#1adil-a2024-07-25 22:50
    HumanEval Benchmark
  • Pull request#2xeon272024-07-25 22:10
  • Issue comment#1jjallaire-aisi2024-07-25 19:56
    HumanEval Benchmark
  • Pull request#1adil-a2024-07-24 19:24

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 0 stars here means stars gained during the window, not the repo's star count.