AgentOps Toolkit is an open source framework and CLI for adding continuous evaluation and observability to enterprise AI agents. It standardizes evaluation patterns, automates assessments in CI/CD workflows, and generates structured signals that help teams monitor, control, and safely operate agentic systems at scale.
active 2026-04-15 → 2026-08-09 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 117 days from 2026-04-15 to 2026-08-09. Pushes: 246 total, peak 38 in a day. Pull requests: 14 total, peak 7 in a day. Issues: 13 total, peak 4 in a day. Comments: 5 total, peak 2 in a day. Stars: 0 total, peak 0 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| placerda | 215 | 193 | 11 | 2 |
| github-actions[bot] | 29 | 29 | 0 | 0 |
| Dongbumlee | 28 | 20 | 1 | 3 |
| dependabot[bot] | 6 | 4 | 2 | 0 |
Recent activity
Latest issues, pull requests and releases
- Issue#396placerda2026-08-07 21:12docs(cicd): explain why the generated safety-eval job has no environment
- Pull request#392placerda2026-08-07 21:08
- Pull request#384placerda2026-08-07 14:46
- Issue#343placerda2026-07-01 22:06Ship Azure Monitor dashboard for Foundry / Azure OpenAI operational metrics
- Issue comment#151placerda2026-06-12 20:53AI-assisted evaluators fail with HTTP 400 when judge deployment is a reasoning model (gpt-5.x, o1, o3, o4)
- Pull request#237placerda2026-06-04 20:42
- Pull request#208placerda2026-05-29 15:31
- Pull request#204placerda2026-05-29 08:19
- Pull request#194placerda2026-05-29 03:44
- Pull request#194placerda2026-05-29 03:44
- Pull request#191placerda2026-05-29 02:26
- Pull request#190placerda2026-05-29 02:18
- Pull request#190placerda2026-05-29 02:17
- Pull request#186Dongbumlee2026-05-28 20:58
- Pull request#166dependabot[bot]2026-05-15 15:48
- Pull request#165dependabot[bot]2026-05-15 15:48
- Issue#88placerda2026-05-08 16:06test: validate eval traces in Azure Monitor (App Insights + GenAI blade)
- Issue#138placerda2026-05-08 15:59Plan AgentOps landing
- Issue#126placerda2026-05-08 15:57Validate tutorial: model-direct evaluation
- Issue#117Dongbumlee2026-04-29 21:49agent_workflow_baseline silently scores ~0 against tool-less agents
- Issue#116Dongbumlee2026-04-29 21:49Default avg_latency_seconds threshold (10s) is unrealistic for Foundry cloud eval
- Issue comment#115Dongbumlee2026-04-29 21:49Implement or hide 'agentops config validate' (currently a stub)
- Issue comment#115Dongbumlee2026-04-29 21:25Implement or hide 'agentops config validate' (currently a stub)
- Issue#114Dongbumlee2026-04-29 21:18results.json schema_version is missing/None
- Issue#116Dongbumlee2026-04-29 21:18Default avg_latency_seconds threshold (10s) is unrealistic for Foundry cloud eval
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 0 stars here means stars gained during the window, not the repo's star count.