Skip to content
View hallieren's full-sized avatar

Block or report hallieren

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. ai-agent-evaluation ai-agent-evaluation Public

    AI Agent Evaluation: score the endpoint, attribute the path, account for the side effects. A free open-source book on evaluating LLM agents, with templates and zero-dependency labs.

    Python

  2. the-last-mile the-last-mile Public

    The Last Mile: a demo is L0, delivery is L4. A free open-source field guide to taking enterprise AI systems from demo to daily use, with 24 field templates and zero-dependency scripts.

    Python

  3. research-rewritten research-rewritten Public

    Research, Rewritten: using AI to produce knowledge you can trust. A free open-source book on AI-assisted research, with prompts, templates, and two rerunnable preregistered experiments.

    Python

  4. substack-data substack-data Public

    Data, scripts, and full agent trajectories behind the AI Agent Evaluation Substack series: a sealed SWE-bench Verified audit, eval sets, judge calibration, red-team rounds, online eval, and release…

    Python

  5. pico pico Public

    A minimal agent harness with production teeth: streaming, parallel tool calls, subagents, compaction, approval gates, spend ceilings, retries, and prompt caching in under 1,000 lines. 78.6% on seal…

    Python