Pinned Loading
-
ai-agent-evaluation
ai-agent-evaluation PublicAI Agent Evaluation: score the endpoint, attribute the path, account for the side effects. A free open-source book on evaluating LLM agents, with templates and zero-dependency labs.
Python
-
the-last-mile
the-last-mile PublicThe Last Mile: a demo is L0, delivery is L4. A free open-source field guide to taking enterprise AI systems from demo to daily use, with 24 field templates and zero-dependency scripts.
Python
-
research-rewritten
research-rewritten PublicResearch, Rewritten: using AI to produce knowledge you can trust. A free open-source book on AI-assisted research, with prompts, templates, and two rerunnable preregistered experiments.
Python
-
substack-data
substack-data PublicData, scripts, and full agent trajectories behind the AI Agent Evaluation Substack series: a sealed SWE-bench Verified audit, eval sets, judge calibration, red-team rounds, online eval, and release…
Python
-
pico
pico PublicA minimal agent harness with production teeth: streaming, parallel tool calls, subagents, compaction, approval gates, spend ceilings, retries, and prompt caching in under 1,000 lines. 78.6% on seal…
Python
If the problem persists, check the GitHub status page or contact support.