Beta-stage, TypeScript-first OpenAI-compatible orchestration runtime for Node and Bun.
Use native when you want exact SDK behavior.
Use agent when you want protocol-guided orchestration, tools, subagents, retries, context management, provider key rotation, request queueing, and cleaned output.
General.AI is not a thin wrapper. It is a protocol-guided orchestration runtime designed to make model behavior more stable and controllable.
Tested heavily on NVIDIA-compatible OpenAI-style endpoints. Broader provider validation is in progress.
This README follows the current beta track of General.AI. If you are on the stable
latestchannel, newer capabilities such as context management/compression, structured checkpoints, parallel action batching, provider key rotation, provider request queueing,classic_v2compatibility, runtime presets, intelligence-aware prompt guidance, and soft-requireddonehandling may not be available yet. Use the beta install instructions below when you want the features called out in the Beta Changelog.
- npm: https://npmjs.com/package/@lightining/general.ai
- GitHub: https://github.com/nixaut-codelabs/general.ai
- Why General.AI
- Why Use It
- Install
- Beta Install
- Stable And Beta
- How It Compares
- Quick Start
- Killer Demo
- Native And Agent
- Compatibility Profiles
- Presets And Intelligence
- Provider Pools And Queue
- Tools
- Subagents
- Thinking, Safety, And Context
- Observability
- Prompt Overrides
- Streaming
- Testing
- Beta Track Highlights
- Package Notes
- License
Most projects end up in one of two bad places:
- they stay very close to the raw provider API and rebuild orchestration from scratch
- or they use a wrapper that hides too much and makes advanced provider behavior harder to reach
General.AI tries to sit in the middle:
nativekeeps the OpenAI client shape intactagentadds a controllable orchestration runtime on top
That means you can stay close to the transport layer when you want, and move up to a higher-level runtime when you need more stability, structure, and visibility.
Use General.AI when you want:
- more stable behavior from smaller or inconsistent models
- a protocol-guided runtime instead of ad hoc prompt glue
- tools, subagents, retries, cleaned output, and context handling in one place
- provider-managed round-robin key rotation and request queueing for OpenAI-compatible gateways
- visibility into why the runtime called a tool, opened a subagent, or compacted context
- direct access to OpenAI-compatible APIs without losing provider-native escape hatches
Do not use it if all you want is a very thin helper around the OpenAI SDK. In that case, stay on native.
npm install @lightining/general.ai openaior:
bun add @lightining/general.ai openaiRuntime targets:
- Node
>=22 - Bun
>=1.1.0
General.AI is ESM-only.
If you want the current beta track:
npm install @lightining/general.ai@beta openaior:
bun add @lightining/general.ai@beta openaiChannel guide:
latest: slower-moving stable channelbeta: newest runtime features, compatibility work, and beta-only capabilities documented in this README
If you only want the stable channel, stay on latest.
Current beta target:
1.2.0-beta.1
General.AI now has two channels on purpose.
| Channel | Recommended when | What to expect | Tradeoff |
|---|---|---|---|
latest |
You want the slower-moving release line. | Smaller, steadier surface area and fewer moving parts. | Newer runtime capabilities can take longer to land. |
beta |
You want the newest orchestration/runtime work. | Faster iteration on context management, compatibility work, provider pools, queueing, presets, intelligence guidance, and parser/recovery improvements. | More behavior may still be refined before it graduates to latest. |
Beta-only or beta-track capabilities documented in this README:
| Capability | latest |
beta |
|---|---|---|
| Provider-managed API key rotation and request queueing | Do not assume | Yes |
| Context summarize / drop strategies | Do not assume | Yes |
classic_v2 compatibility profile |
Do not assume | Yes |
Runtime presets and intelligence guidance |
Do not assume | Yes |
Soft-required done with inferred completion |
Do not assume | Yes |
| Built-in speed metrics and stream TPS reporting | Do not assume | Yes |
| Ongoing parser/recovery hardening | Slower-moving | Faster-moving |
If you need the features above today, install @lightining/general.ai@beta.
General.AI is intentionally narrower than big frameworks. That is part of the pitch, not a bug.
| Library | Best when | Where General.AI is stronger | Where General.AI is weaker |
|---|---|---|---|
| Raw OpenAI SDK | You want the official API surface with minimal abstraction. | Adds protocol-guided agents, cleaned output, tool/subagent orchestration, context controls, provider key rotation, and request queueing on top of an OpenAI-compatible transport. | Much smaller surface, smaller ecosystem, and not an official SDK. If you only need direct API access, the raw SDK is simpler. |
| LangChain | You want a large integration ecosystem, prebuilt agent abstractions, and the broader LangChain/LangGraph/LangSmith stack. | More explicit about low-level protocol behavior, provider shaping, prompt assembly, cleaned-vs-raw output, and provider operations like queueing and rate-limit handoff. | Far fewer integrations, much smaller ecosystem, less mature tracing/deployment tooling, and less battle-tested overall. |
| Vercel AI SDK | You want a unified provider API plus strong frontend/UI hooks and streaming patterns. | More focused on backend orchestration internals, protocol parsing, subagent/tool loops, and runtime recovery behavior. | Not a UI toolkit, not framework-first, and much smaller in provider coverage and frontend ergonomics. |
The honest shorthand:
- choose General.AI when you want an OpenAI-compatible orchestration runtime with explicit control over the runtime loop
- choose the raw OpenAI SDK when you want official transport access and almost no abstraction
- choose LangChain when you want breadth, integrations, and a larger agent ecosystem
- choose the AI SDK when you want a unified provider layer with strong TypeScript and frontend ergonomics
import OpenAI from "openai";
import { GeneralAI } from "@lightining/general.ai";
const openai = new OpenAI({
apiKey: process.env.OPENAI_API_KEY,
});
const generalAI = new GeneralAI({ openai });
const result = await generalAI.agent.generate({
endpoint: "chat_completions",
model: "gpt-5.4-mini",
messages: [
{ role: "user", content: "Say hello briefly in Turkish." },
],
});
console.log(result.cleaned);import OpenAI from "openai";
import { GeneralAI } from "@lightining/general.ai";
const openai = new OpenAI({
apiKey: process.env.OPENAI_API_KEY,
});
const generalAI = new GeneralAI({ openai });
const result = await generalAI.agent.generate({
endpoint: "chat_completions",
model: "gpt-5.4-mini",
preset: "classic_safe",
intelligence: "high",
messages: [
{ role: "user", content: "Say hello briefly in Turkish." },
],
});
console.log(result.cleaned);
console.log(result.meta.warnings);
console.log(result.meta.performance.speed);
console.log(result.meta.configuration);You can also let General.AI construct and manage the provider client for you:
import { GeneralAI } from "@lightining/general.ai";
const generalAI = new GeneralAI({
provider: {
baseURL: "https://integrate.api.nvidia.com/v1",
apiKeys: [
process.env.NVIDIA_KEY_A!,
process.env.NVIDIA_KEY_B!,
],
rotation: {
strategy: "round_robin",
onRateLimit: "next_key",
maxRateLimitHandoffs: 2,
},
queue: {
maxConcurrentRequests: 3,
},
},
});Returned shape:
type GeneralAIAgentResult = {
output: string;
cleaned: string;
events: ProtocolEvent[];
meta: {
warnings: string[];
prompt: RenderedPrompts;
strippedRequestKeys: string[];
configuration: {
preset:
| "balanced"
| "strict"
| "fast"
| "agentic"
| "classic_safe"
| "research";
intelligence: "minimal" | "medium" | "high";
compatibilityProfile: "auto" | "modern" | "classic" | "classic_v2";
safetyEnabled: boolean;
thinkingEnabled: boolean;
};
completion: {
explicitDone: boolean;
inferredDone: boolean;
};
stepCount: number;
toolCallCount: number;
subagentCallCount: number;
protocolErrorCount: number;
contextOperations: string[];
contextSummaryCount: number;
contextDropCount: number;
memorySessionId?: string;
performance: {
wallTimeMs: number;
requestTimeMs: number;
timeToFirstTokenMs?: number;
speed: {
mode: "heuristic_speed_index" | "stream_tps";
unit: "speed_index" | "tokens_per_second";
value: number;
label: "very_slow" | "slow" | "steady" | "fast" | "very_fast";
algorithm: string;
};
steps: Array<{
step: number;
stream: boolean;
durationMs: number;
firstTokenLatencyMs?: number;
outputWindowMs?: number;
}>;
};
endpointResults: unknown[];
};
usage: {
inputTokens: number;
outputTokens: number;
totalTokens: number;
cachedInputTokens: number;
reasoningTokens: number;
};
endpointResult: unknown;
};This is the kind of call where General.AI starts to feel different from a thin wrapper:
const result = await generalAI.agent.generate({
endpoint: "chat_completions",
model: "gpt-5.4-mini",
messages: [
{
role: "user",
content: "Use tools if needed, delegate arithmetic to a subagent if useful, and give me a short final answer.",
},
],
compatibility: {
profile: "classic_v2",
},
tools: {
registry: [weatherTool, calculatorTool],
},
subagents: {
registry: [mathHelper],
},
context: {
mode: "auto",
strategy: "hybrid",
},
});
console.log(result.cleaned);
console.log(result.meta.contextOperations);
console.log(result.meta.warnings);In one runtime call, General.AI can:
- call one or more tools
- delegate to one or more subagents
- retry after malformed protocol output
- summarize or drop older context
- return cleaned user-visible output separately from raw protocol output
General.AI exposes two surfaces.
Use native when you want exact OpenAI SDK behavior.
const response = await generalAI.native.responses.create({
model: "gpt-5.4-mini",
input: "Give a one-sentence explanation of prompt caching.",
});
const completion = await generalAI.native.chat.completions.create({
model: "gpt-5.4-mini",
messages: [
{ role: "user", content: "Say hello in one sentence." },
],
});This keeps:
- request bodies OpenAI-native
- response objects OpenAI-native
- stream events OpenAI-native
- optional provider-managed key rotation and request queueing when you construct from
provider
Use agent when you want runtime orchestration.
The agent runtime can:
- assemble layered prompts
- enforce a structured text protocol
- parse runtime events from model output
- retry recoverable protocol failures
- call tools and subagents
- maintain optional memory
- summarize or drop old context before the model hits its limit
- return both raw protocol output and cleaned user-visible output
Some OpenAI-compatible providers are stricter than others about message roles and continuation shaping.
General.AI supports compatibility profiles:
modernclassicclassic_v2auto
Example:
compatibility: {
profile: "classic_v2",
}What they mean:
modern: modern OpenAI-style behaviorclassic: safer classicsystem/user/assistantshapingclassic_v2: stricter provider-safe continuation shaping for gateways that dislike late system-style messagesauto: currently resolves tomodernunless explicitly overridden
If you are using stricter compatible gateways, classic_v2 is the safest place to start.
General.AI now separates provider shaping from model guidance intensity.
Use compatibility.profile for provider behavior, and use preset plus intelligence for runtime posture.
Available presets:
balancedstrictfastagenticclassic_saferesearch
Available intelligence levels:
minimalmediumhigh
Example:
const result = await generalAI.agent.generate({
endpoint: "chat_completions",
model: "gpt-5.4-mini",
preset: "agentic",
intelligence: "high",
messages: [{ role: "user", content: "Solve this carefully." }],
});What they mean:
preset: chooses a higher-level runtime posture such asfast,strict, orresearchintelligence: changes how heavy-handed the prompt guidance isminimal: more explicit protocol guidance for weaker or less consistent modelsmedium: balanced default guidancehigh: lighter protocol guidance for stronger models that do not need to be over-instructed
Subagents can also override preset and intelligence.
General.AI beta can manage an OpenAI-compatible provider for you instead of requiring a prebuilt client.
const generalAI = new GeneralAI({
provider: {
name: "nvidia",
baseURL: "https://integrate.api.nvidia.com/v1",
apiKeys: [
{ key: process.env.NVIDIA_KEY_A!, label: "nvidia-a" },
{ key: process.env.NVIDIA_KEY_B!, label: "nvidia-b" },
{ key: process.env.NVIDIA_KEY_C!, label: "nvidia-c" },
],
rotation: {
strategy: "round_robin",
onRateLimit: "next_key",
maxRateLimitHandoffs: 3,
revisitKeysInSameRequest: false,
},
queue: {
enabled: true,
maxConcurrentRequests: 3,
maxQueuedRequests: 100,
strategy: "fifo",
},
},
});Current beta behavior:
- keys are selected in round-robin order
429triggers a handoff to the next unused key for that request- non-
429provider failures are surfaced normally - requests beyond the provider concurrency limit wait in a provider-level FIFO queue
- the same provider queue is shared by root agent calls, subagents, and provider-backed native calls
This is a beta feature and has been validated most heavily on NVIDIA-compatible OpenAI-style endpoints so far.
General.AI tools are runtime-defined JavaScript functions triggered by protocol markers.
import { defineTool } from "@lightining/general.ai";
const echoTool = defineTool({
name: "echo",
description: "Echo text back for runtime testing.",
inputSchema: {
type: "object",
additionalProperties: false,
properties: {
text: { type: "string" },
},
required: ["text"],
},
async execute(args) {
return { echoed: args.text };
},
});Tool access can be scoped:
- root only
- all subagents
- selected subagents only
const rootOnlyTool = defineTool({
name: "root_only",
description: "Only callable from the root agent.",
access: {
subagents: false,
},
async execute() {
return { ok: true };
},
});The runtime also supports multiple tool calls in the same step, with configurable parallel limits.
Subagents are bounded delegated General.AI runs with their own config.
import { defineSubagent } from "@lightining/general.ai";
const mathHelper = defineSubagent({
name: "math_helper",
description: "A precise arithmetic specialist.",
instructions: "Solve delegated arithmetic carefully and return a concise answer.",
model: "gpt-5.4-mini",
request: {
chat_completions: {
temperature: 0.1,
},
},
});Subagents can override:
endpointmodelpresetintelligencerequestpersonalitysafetythinkingcontextpromptslimitstoolssubagentscompatibilitymemory
They can also participate in parallel action batches.
These systems are separate on purpose.
thinking: {
enabled: true,
mode: "hybrid",
strategy: "checkpointed",
checkpointFormat: "structured",
effort: "high",
}Available thinking modes:
noneinlineorchestratedhybrid
safety: {
enabled: true,
mode: "balanced",
input: {
enabled: true,
},
output: {
enabled: true,
},
}Safety runs inside the agent protocol instead of forcing separate moderation-style API calls for every step.
If safety.enabled is false, the safety prompt section and safety marker requirements are omitted from the assembled runtime prompt.
context: {
enabled: true,
mode: "auto",
strategy: "hybrid",
trigger: {
contextRatio: 0.9,
},
}Supported context strategies:
summarizedrop_oldestdrop_nonessentialhybrid
Supported modes:
offautomanualhybrid
This is runtime-managed context control. It is not a built-in provider compression feature.
[[[status:done]]] is still the preferred explicit finalizer, but it is no longer hard-required.
If the model ends on a clearly complete final writing block, the runtime can infer completion and record that in:
result.meta.completionGeneral.AI is designed to be inspectable.
You can already inspect:
- parsed protocol events
- warnings and retry reasons
- cleaned output and raw protocol output
- runtime configuration, including
preset,intelligence, and resolved compatibility profile - completion mode, including whether
donewas explicit or inferred - tool and subagent counts
- prompt rendering output
- context compaction operations
- endpoint result history
- performance timing and speed metrics
- provider-backed queueing and retry warnings
This helps answer questions like:
- why did it call a tool?
- why did it open a subagent?
- why did it summarize or drop old messages?
- why did it retry after malformed model output?
General.AI renders a layered prompt stack in this order:
- identity
- endpoint adapter rules
- protocol
- safety
- personality
- thinking
- tools and subagents
- memory
- task context
Bundled prompts live in prompts/*.txt.
Prompt placeholders:
{data:key}for scalar values{block:key}for multiline blocks
Example:
const prompt = await generalAI.agent.renderPrompts({
endpoint: "responses",
model: "gpt-5.4-mini",
messages: [{ role: "user", content: "Hello" }],
prompts: {
sections: {
task: "Task override.\n{block:task_context}",
},
},
});Raw overrides are also supported:
prompts: {
raw: {
prepend: "Extra preamble",
append: "Extra appendix",
replace: "Replace the entire rendered prompt",
},
}Use native for exact provider stream events.
Use agent.stream() for parsed runtime events and cleaned writing deltas.
const stream = generalAI.agent.stream({
endpoint: "responses",
model: "gpt-5.4-mini",
messages: [{ role: "user", content: "Say hello." }],
});
for await (const event of stream) {
if (event.type === "writing_delta") {
process.stdout.write(event.text);
}
}Common stream events:
run_startedprompt_renderedstep_startedraw_text_deltawriting_deltaprotocol_eventbatch_startedtool_startedtool_resultsubagent_startedsubagent_resultcontext_compactedwarningrun_completed
The streaming path also includes recovery for malformed protocol output from real models.
Deterministic tests:
npm testCross-runtime smoke tests:
npm run smokeManual public-surface walkthrough:
bun run test.jstest.js can also exercise optional live provider checks when environment variables are set. It covers:
- native chat
- agent protocol generation
- parallel tool batching
- subagent delegation
- orchestrated thinking
- context summarization
- context dropping
- streaming
Useful environment variables:
GENERAL_AI_API_KEY=...
GENERAL_AI_BASE_URL=...
GENERAL_AI_MODEL=...
GENERAL_AI_SKIP_LIVE=1If GENERAL_AI_SKIP_LIVE=1 is set, the broader manual scripts skip live provider checks.
Install the beta channel:
npm install @lightining/general.ai@beta openaior:
bun add @lightining/general.ai@beta openaiThis is a channel-level overview, not a historical beta.0-to-beta.1 changelog.
Current beta-track highlights:
- parallel tool and subagent action batching
- runtime presets plus model-capacity-aware
intelligenceguidance - subagent-specific models and endpoint request parameters
- thinking modes:
inline,orchestrated,hybrid - structured
checkpointandrevisesupport - context management with summarize / drop / hybrid strategies
- provider-managed round-robin API key rotation
429handoff to the next key without revisiting exhausted keys in the same request- provider-level FIFO request queueing
- disabled subsystems are omitted from the assembled runtime prompt instead of being described as "off"
- soft-required
donewith inferred completion when the final writing block is clearly complete - built-in per-run speed metrics with heuristic speed indexing and stream TPS reporting
- stronger streaming fallback and retry behavior
- parser recovery for plain-text safety-style payload blocks
- duplicate final-writing guards that stop obvious repeat loops earlier
- compatibility profiles including
classic_v2
Features listed above are beta-track features. If you install @lightining/general.ai without @beta, you may be on an older stable release that does not include all of them yet.
Beta reality check:
- protocol compliance still depends on model quality
- some providers are stricter than others about message shaping
- broader provider validation is still in progress
Bundled prompts are written in English for consistency, but user-visible output still follows the user’s language unless they explicitly ask for another one.
General.AI is ESM-only.
The current SDK baseline is openai@^6.33.0.
General.AI beta is aimed at:
- app backends
- internal LLM runtimes
- tool and subagent orchestration layers
- OpenAI-compatible provider integrations
It is not intended as a browser bundle.
Apache-2.0. See LICENSE.