Minimal, provider-agnostic LLM client for Go. Streaming completions, tool calling, prompt caching, token counting, cost projection, and retry middleware for Anthropic Claude (Messages API), OpenAI (Chat Completions + Responses API, including GPT-5 family), Google Gemini (text / image / video), and any OpenAI-compatible endpoint (Azure OpenAI, Groq, Together, vLLM, OpenRouter, Ollama). Idiomatic Go: iter.Seq2 streaming, sealed sum types, errors.Is sentinels, context.Context cancellation.
Status: v1.0.0 — stable. The public API follows semver: no breaking changes without a major (module-path) bump. See CHANGELOG.md. Dogfooded in production at Noumenal.
go get github.com/amit-timalsina/pi-llm-goRequires Go 1.23+ (for iter.Seq2).
| Capability | Anthropic | OpenAI Chat | OpenAI Responses | Gemini |
|---|---|---|---|---|
| Streaming text | ✅ | ✅ | ✅ | ✅ |
| Tool calling | ✅ | ✅ | ✅ | ✅ |
Forced tool choice (Request.ToolChoice) |
✅ | ✅ | ✅ | ✅ |
Strict / grammar-constrained tool input (Tool.Strict) |
✅ | ✅ | ✅ | ✅ |
| Image input | ✅ | ✅ | ✅ | ✅ |
| Video input | reject at wire | reject at wire | reject at wire | ✅ native |
| Extended thinking | ✅ | — | ✅ (reasoning summaries) | ✅ |
| Prompt caching (5m + 1h tier) | ✅ | automatic | automatic | single-TTL |
| Per-TTL cache-write Usage breakdown | ✅ | — | — | — |
CountTokens (no-spend pre-flight) |
✅ | — (no API) | — (no API) | ✅ |
| Cost projection helper | ✅ | ✅ | ✅ | ✅ |
Retry middleware (Options.Retry) |
✅ | ✅ | ✅ | ✅ |
| Files API helper (>20 MB inputs) | — | — | — | ✅ |
package main
import (
"context"
"fmt"
"os"
llm "github.com/amit-timalsina/pi-llm-go"
"github.com/amit-timalsina/pi-llm-go/providers/anthropic"
)
func main() {
p, _ := anthropic.New(anthropic.Options{APIKey: os.Getenv("ANTHROPIC_API_KEY")})
msg, err := llm.Complete(context.Background(), p, llm.Request{
Model: anthropic.ClaudeSonnet4_6,
MaxTokens: 1024,
Messages: []llm.Message{
{Role: llm.RoleUser, Content: []llm.Block{llm.TextBlock{Text: "Reply with one word: hello"}}},
},
})
if err != nil { panic(err) }
for _, block := range msg.Content {
if tb, ok := block.(llm.TextBlock); ok {
fmt.Println(tb.Text)
}
}
}for event, err := range p.Stream(ctx, req) {
if err != nil { /* handle */ break }
if d, ok := event.(llm.EventTextDelta); ok {
fmt.Print(d.Delta)
}
}Stream returns iter.Seq2[llm.StreamEvent, error]. Complete is a synchronous helper that drains the stream and returns the final assistant Message.
req := llm.Request{
Model: anthropic.ClaudeSonnet4_6,
MaxTokens: 1024,
Tools: []llm.Tool{{
Name: "get_weather",
Description: "Get the weather for a city",
InputSchema: json.RawMessage(`{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}`),
}},
Messages: []llm.Message{
{Role: llm.RoleUser, Content: []llm.Block{llm.TextBlock{Text: "What's the weather in Tokyo?"}}},
},
}
msg, _ := llm.Complete(ctx, p, req)
for _, b := range msg.Content {
if tc, ok := b.(llm.ToolCallBlock); ok {
fmt.Printf("tool=%s args=%s\n", tc.Name, tc.Arguments)
// Execute the tool yourself, then send a ToolResultBlock on the next turn.
// For a built-in execution loop, use https://github.com/amit-timalsina/pi-agent-go
}
}Force a tool call with Request.ToolChoice — Auto (default), Any (must call some tool), Tool (must call a named tool), or None (text only):
req.ToolChoice = llm.ToolChoice{Type: llm.ToolChoiceTool, Name: "get_weather"}Constrain arguments to the schema with Tool.Strict: true. The provider then samples tool arguments under a grammar derived from InputSchema, so enum fields and required keys come back exactly as declared instead of being silently dropped — no caller-side validate-and-retry loop. Wired across all four providers (Anthropic, OpenAI Chat, OpenAI Responses, Gemini):
Tools: []llm.Tool{{Name: "set_valve", InputSchema: schema, Strict: true}}See examples/strict_tool_use for the end-to-end pattern.
| You want | Pick |
|---|---|
| One Go interface across Anthropic / OpenAI / Gemini, streaming-first | pi-llm-go |
| Only OpenAI, no streaming abstraction needed | sashabaranov/go-openai |
| Only Anthropic, vendor SDK guarantees | anthropics/anthropic-sdk-go |
| A heavyweight framework with chains, agents, memory, embeddings, vector stores | tmc/langchaingo |
| Built-in agent loop on top of a provider-agnostic LLM client | pi-llm-go + pi-agent-go |
pi-llm-go is the streaming-completions layer. ~1.5kLoC of plain Go: one interface, four providers, no model registries, no plugin systems, no abstractions beyond what the LLM wire format demands.
- Streaming-first.
Stream()returnsiter.Seq2[StreamEvent, error]— Go 1.23 iterators, no callbacks, no goroutine leaks.Complete()is the synchronous helper for one-shot use. - Sealed sum types.
BlockandStreamEventare interfaces with package-private marker methods. Type-switch exhaustively; the compiler tells you if you miss a case. - Tool calling. Declare tools on
Request.Tools; receiveToolCallBlocks on the response; sendToolResultBlocks back. Pi-llm-go does not execute tools — that's pi-agent-go's job. - Extended thinking. Set
Request.Thinkingand readThinkingBlockcontent back. Two shapes on Anthropic: adaptiveThinkingConfig{Effort: llm.EffortMedium}(Opus 4.6+, required on 4.7+) or legacyThinkingConfig{BudgetTokens: N}(Sonnet/Haiku);Effortwins when both are set. AddDisplay: llm.DisplaySummarizedto the adaptive shape to stream summarized thinking text — on Opus 4.7+ the default is signature-only blocks with no text. OpenAI Responses exposes reasoning via its ownOptions.ReasoningEffort; Gemini thinks natively. - Open-closed providers. Implement
LLM.Streamto add custom providers; no plugin registry needed. - Errors that branch cleanly.
errors.Is(err, llm.ErrRateLimit)works through*APIErrorwraps;errors.As(err, &apiErr)gives you status + body. - Cancellation =
context.Context. No bespoke abort signal types.
import "github.com/amit-timalsina/pi-llm-go/providers/anthropic"
p, _ := anthropic.New(anthropic.Options{
APIKey: os.Getenv("ANTHROPIC_API_KEY"),
// BaseURL: defaults to https://api.anthropic.com
// Version: defaults to "2023-06-01"
// Beta: optional anthropic-beta header values
})Honors Request.Thinking. Surfaces all content-block types (text, thinking, tool_use).
import "github.com/amit-timalsina/pi-llm-go/providers/openai"
p, _ := openai.New(openai.Options{
APIKey: os.Getenv("OPENAI_API_KEY"),
BaseURL: "https://api.openai.com/v1", // or "https://api.groq.com/openai/v1", etc.
})Talks the Chat Completions wire format, so the same provider works against any compatible host. Request.Thinking is ignored at v1 — reasoning-effort dialects vary too much across compatible hosts to map portably.
import "github.com/amit-timalsina/pi-llm-go/providers/gemini"
p, _ := gemini.New(gemini.Options{APIKey: os.Getenv("GEMINI_API_KEY")})Native support for the Gemini 2.5 / 3 / Robotics ER 1.6 families. Same LLM interface as the other providers, plus llm.VideoBlock for native video input (Gemini is the only provider that accepts video natively; Anthropic and OpenAI reject VideoBlock at the wire boundary with a clear pointer to the frame-extraction workaround). YouTube URLs work directly:
llm.VideoBlock{URI: "https://www.youtube.com/watch?v=..."}For files larger than ~20 MB, the sibling providers/gemini/files sub-package handles the multipart upload + ACTIVE-state polling:
import "github.com/amit-timalsina/pi-llm-go/providers/gemini/files"
fc, _ := files.New(files.Options{APIKey: os.Getenv("GEMINI_API_KEY")})
ref, _ := fc.Upload(ctx, mp4Reader, "video/mp4", files.UploadOptions{DisplayName: "demo.mp4"})
ref, _ = fc.Wait(ctx, ref, files.WaitOptions{}) // polls until ACTIVE
defer fc.Delete(context.Background(), ref.Name) // ~48h server-side TTL if you forget
// Now use the URI in a generateContent call:
content := []llm.Block{llm.TextBlock{Text: "describe"}, llm.VideoBlock{URI: ref.URI}}Vertex AI (gs:// URIs + OAuth) is a planned future addition; the Gemini provider currently targets the Google AI direct endpoint.
Non-2xx HTTP responses surface as *llm.APIError wrapping one of the typed sentinels — errors.Is works through the wrapping so consumers branch on category, not status code:
ErrAuth // 401, 403
ErrRateLimit // 429
ErrInvalidRequest // other 4xx (parent of the next two)
├─ ErrContextLength // prompt / max_tokens exceeds the model's window
└─ ErrPolicyViolation // input rejected by content / safety policy
ErrProvider // generic provider problem (parent of the next two)
├─ ErrServerError // 5xx (excluding 529)
└─ ErrOverloaded // 529 (Anthropic infra overload)
ErrServerError and ErrOverloaded both wrap ErrProvider; ErrContextLength and ErrPolicyViolation both wrap ErrInvalidRequest. Legacy errors.Is(err, ErrProvider) and errors.Is(err, ErrInvalidRequest) callers keep matching the full subtree.
The finer 4xx sentinels are derived by inspecting the response body — provider error schemas don't carry a canonical machine-readable category for "context too long" vs "policy violation," so pi-llm-go pattern-matches the message text. Patterns cover Anthropic / OpenAI / Gemini current shapes. For structured per-provider decoding, branch on apiErr.Body directly.
Sugar helpers and the parsed Retry-After:
for ev, err := range provider.Stream(ctx, req) {
if err == nil { /* consume ev */ continue }
var apiErr *llm.APIError
if errors.As(err, &apiErr) {
switch {
case llm.IsRateLimited(err): // 429 → respect apiErr.RetryAfter
case llm.IsOverloaded(err): // 529 → short backoff, consider failover
case llm.IsServerError(err): // 5xx → retry, escalate if sustained
case errors.Is(err, llm.ErrAuth):
}
}
return err
}APIError.RetryAfter is populated by all four built-in providers when the response carries a Retry-After (RFC 7231 delta-seconds or HTTP-date) or retry-after-ms (OpenAI's millisecond form, which wins when both are present).
Set Options.Retry on any provider to retry retriable errors (429 / 529 / 5xx / transient network failures) with exponential backoff + full jitter:
p, _ := anthropic.New(anthropic.Options{
APIKey: key,
Retry: &llm.RetryPolicy{MaxAttempts: 4, BaseDelay: time.Second, MaxDelay: 30 * time.Second},
})
// or use the sane defaults:
p2, _ := anthropic.New(anthropic.Options{APIKey: key, Retry: ptr(llm.DefaultRetryPolicy())})Server-supplied Retry-After hints dominate the exponential schedule (capped at MaxDelay). ErrAuth, ErrInvalidRequest, and context.Canceled/DeadlineExceeded are NOT retried — caller intent always wins.
Scope: retry covers only the initial HTTP attempt. Once a 200 OK lands and streaming begins, the run is committed — a mid-stream connection break terminates the iterator rather than retrying, because resuming would replay events the consumer already saw. Callers needing at-least-once streaming should wrap pi-llm-go in their own idempotent replay layer.
llm.RunWithRetry[T] is exported so callers can build cross-provider fallback / circuit-breaker logic on top.
Telemetry: every retry emits a slog.DebugContext record so silent backoff windows are visible to ops:
DEBUG llm.retry.attempt attempt=1 max_attempts=4 delay_ms=820 cause=overloaded retry_after_ms=500 error="..."
DEBUG llm.retry.exhausted max_attempts=4 last_cause=overloaded error="..."
Default slog level is INFO so these are silent unless you configure DEBUG. Field names are part of the v1 contract — write a custom slog.Handler to route them to Prometheus counters, dashboards, or wherever else. Otel users can bridge via otelslog.
Every completion surfaces a typed Usage value on the final EventMessageEnd (and on the *llm.Message returned by Complete / Accumulate):
msg, _ := llm.Complete(ctx, provider, req)
fmt.Printf("in=%d out=%d cache_write=%d (5m=%d, 1h=%d) cache_read=%d total=%d\n",
msg.Usage.InputTokens, msg.Usage.OutputTokens,
msg.Usage.CacheWriteTokens, msg.Usage.CacheWrite5mTokens, msg.Usage.CacheWrite1hTokens,
msg.Usage.CacheReadTokens, msg.Usage.TotalTokens)CacheWrite5mTokens and CacheWrite1hTokens are Anthropic-specific — they break CacheWriteTokens down by TTL tier so consumers can detect silent 5min fallback when CacheRetention=long was requested but the model didn't honor the extended TTL:
if req.CacheRetention == llm.CacheRetentionLong &&
msg.Usage.CacheWrite5mTokens > 0 && msg.Usage.CacheWrite1hTokens == 0 {
// Model downgraded the cache hold to 5min. Cost projections that
// assumed a 1h-cached prefix across iterations need to adjust.
}OpenAI and Gemini leave the TTL-breakdown fields at zero (their cache surfaces are opaque or single-TTL).
llm.ComputeCost applies a built-in pricing table to a Usage value and returns a dollar breakdown:
cost, err := llm.ComputeCost(msg.Usage, req.Model)
if err == nil {
fmt.Printf("$%.4f total (in=$%.4f out=$%.4f cache_read=$%.4f cache_write_1h=$%.4f)\n",
cost.Total(), cost.Input, cost.Output, cost.CacheRead, cost.CacheWrite1h)
}The seed table covers the Claude 4, GPT-5, and Gemini 2.5/3.1 families (verified 2026-05-13). For any model not in the table — older snapshots, deprecated IDs, regional/batch tiers — call llm.RegisterPricing(modelID, llm.Pricing{...}) at startup with rates pulled from your own source. Registered entries override the seed.
The TTL-breakdown fields from Usage carry through automatically: when CacheWrite5mTokens and CacheWrite1hTokens are populated, they price against their respective tiers independently — so silent 5min fallback is reflected in the cost projection without any caller-side branching.
Anthropic and Gemini expose a count-tokens endpoint that returns the input-token count without spending an inference call. Providers implement llm.TokenCounter when they support this:
if c, ok := provider.(llm.TokenCounter); ok {
n, err := c.CountTokens(ctx, req)
if err == nil {
fmt.Printf("request would consume %d input tokens\n", n)
}
}OpenAI's Chat Completions and Responses providers do NOT implement TokenCounter — OpenAI's tokenization is local-only via tiktoken, which pi-llm-go does not bundle. Callers needing a pre-flight count for an OpenAI-hosted model should run tiktoken themselves.
Runnable examples in examples/:
examples/streaming— basic streaming of a text response.examples/tool_calling— hand-rolled tool-call loop againstget_current_time.examples/strict_tool_use—Tool.Strict+Request.ToolChoiceforcing schema-valid arguments.examples/multimodal— image input on Anthropic + OpenAI.examples/multimodal_gemini— text / image / video against Gemini;--video-upload PATHexercises the Files API end-to-end.examples/prompt_caching—CacheRetentionknob driving Anthropic prompt-cache hits.examples/thinking— extended thinking (Effort/BudgetTokens) withThinkingBlockoutput.examples/openai_responses— OpenAI Responses API with reasoning summaries.examples/azure_openai— OpenAI-compatible provider against an Azure OpenAI deployment.examples/multi_turn— multi-turn conversation accumulating assistant turns.
Run them with go run ./examples/streaming (set ANTHROPIC_API_KEY first).
As of v1.0.0 this package follows semver strictly: the exported API is stable, and no breaking change ships without a major-version (module-path) bump per Go's major-version policy. New providers, new optional Request/Options fields, and new sealed Block/StreamEvent variants are minor releases. See CHANGELOG.md for each release.
MIT. See LICENSE.
Designed and named after pi-ai by Mario Zechner (TypeScript, MIT). The wire-level vocabulary and event types follow the upstream's lead; the Go-native API surface (interface-based providers, iterator streaming, sealed sum types) is a from-scratch redesign for Go idioms.
Built and maintained by Amit Timalsina with Claude Code assistance — all design decisions and release calls are human-owned.