Golem: A Type-Safe AI Agent Framework for Go
Golem is an open-source AI agent framework for Go with generics, zero dependencies, MCP support, and built-in retries, budgets, and human approval. Here is why I built it and how it works.
If you search for an AI agent framework in Go, the results are thin. You find a few thin wrappers around provider SDKs, some ports of Python libraries, and a lot of tutorials that end with "now call the API in a loop." That is fine for a demo. It stops being fine the first time an agent has to run inside a real service, talk to a real database, and cost real money.
I build backends in Go, and I wanted an agent to be a normal Go value inside the service instead of a separate Python process reached over a queue. So I wrote Golem, an open-source agent framework for Go. This post covers what it is, the problem it targets, how a run works, and the decisions I made along the way.
The problem with agents in production
An agent is a loop. The model reads a conversation, decides to call a tool, you run the tool, you feed the result back, and you repeat until the model produces an answer. Writing that loop takes an afternoon. Making it dependable takes much longer, because every step can fail in its own way:
- The provider returns a 429 or a 503 halfway through a run.
- The model calls a tool with arguments that do not match your schema.
- The model returns JSON that is almost the shape you asked for.
- A tool needs a human to approve it before anything irreversible happens.
- A loop runs longer than anyone expected and the bill shows up the next morning.
- A run dies on turn six and you lose everything the first five turns produced.
Python has good answers to several of these. Pydantic AI in particular showed that an agent can have typed dependencies, validated output, and a clean API. I used it as a reference while designing Golem. But copying Python mechanisms into Go would have produced something un-Go-like: decorators become awkward, runtime validation hides what the compiler could check, and exceptions as control flow do not fit a language that returns errors.
So the design question was not "how do I port this?" It was "what should an agent look like if I start from Go's strengths?"
What Golem is
Golem is a Go library for building AI agents. The short version of the feature list:
- Agents are generic:
Agent[Deps, Output]. Dependencies and final output are real types. - The module has no external dependencies. It is built on the standard library only.
- It has adapters for OpenAI, Anthropic, Google Gemini, AWS Bedrock, Azure OpenAI, and local models through Ollama or LM Studio. For local and custom endpoints the OpenAI adapter no longer needs an API key (v0.8.5 fixed that).
- It includes an MCP client over stdio and streamable HTTP.
- It ships retries, fallback models, self-correction, token and cost limits, streaming, run events, and human-in-the-loop pauses.
- It includes a few common tools: PDF extraction, Word/Excel/PowerPoint extraction, web fetch, file read, shell, and Agent Skills loading.
- It has a deterministic fake model so you can test agents offline.
It is at v0.8.5 as I write this, with a public API that only changes additively on the way to v1.0.
Install it with:
go get github.com/abubakarsiddik31/golem
The layering is deliberately plain. Your service owns the dependencies, the core owns the run, and everything that touches the outside world sits behind one of two small interfaces.
Why Go for agents at all
Agent work is mostly waiting on a network. People reasonably ask whether language matters. I think it does, for reasons that show up only once the agent leaves the notebook.
An agent service is a long-running process that holds many conversations at once, calls tools concurrently, streams tokens to clients, and has to shut down cleanly. Go's goroutines, context.Context, and single static binary fit that shape well. A deployed agent can be one small container image with no interpreter, no virtualenv, and no dependency tree to audit.
The second reason is the compiler. When a tool depends on a database handle and a tenant ID, I would like the compiler to tell me when I pass the wrong one. In Golem that is a type parameter, not a convention.
A first agent
This is the smallest working agent, taken from the README:
client, err := openai.New(openai.Config{
APIKey: os.Getenv("OPENAI_API_KEY"),
Model: "gpt-4o-mini",
})
if err != nil {
panic(err)
}
agent, err := golem.New[struct{}, string](client,
golem.DecodeFunc[string](func(_ context.Context, r model.Response) (string, error) {
return r.Message.Content, nil
}),
)
if err != nil {
panic(err)
}
result, err := agent.Run(context.Background(), golem.RunContext[struct{}]{}, "Why is Go ideal for AI agents?")
fmt.Println(result.Output)
fmt.Printf("Tokens: %d input, %d output\n", result.Usage.InputTokens, result.Usage.OutputTokens)
Two things are worth noticing. The agent is built once with golem.New and run many times. And it needs two collaborators: a model and an output decoder. The decoder is the line where untrusted model text becomes a value your program is willing to use. Everything else, including instructions, tools, budgets, and schemas, is an option function.
The result carries the typed output, the full normalized conversation, and cumulative token usage. Nothing is thrown away.
Typed tools and dependency injection
Most agent bugs I have seen live in tools. A tool needs a database, an HTTP client, a user identity. Frameworks usually solve this with globals or closures, which makes tools hard to test and easy to leak state through.
In Golem, a tool is a typed value, and its Exec function receives the run's dependency value directly:
type Database struct {
Users map[int]string
}
getUser := tool.MustNew(tool.Tool[Database]{
Name: "get_user",
Description: "Look up a user name by their ID.",
Schema: json.RawMessage(`{
"type": "object",
"properties": {"id": {"type": "integer"}},
"required": ["id"]
}`),
Exec: func(ctx context.Context, db Database, args json.RawMessage) (tool.Result, error) {
var input struct {
ID int `json:"id"`
}
if err := json.Unmarshal(args, &input); err != nil {
return tool.Result{}, err
}
name, ok := db.Users[input.ID]
if !ok {
return tool.Text("User not found"), nil
}
return tool.Text(name), nil
},
})
agent, err := golem.New[Database, string](client, decoder,
golem.WithTools[Database, string](getUser),
)
result, err := agent.Run(ctx, golem.RunContext[Database]{Deps: Database{Users: map[int]string{42: "Alice"}}}, "Who is user 42?")
The schema is written by hand and arguments arrive as raw JSON. That is deliberate. Golem does not use reflection to infer schemas or decode arguments, because arguments are untrusted model output and I would rather you see the validation than have it happen somewhere you cannot read. It costs a few more lines per tool. In return, a tool's behavior is visible in the file where it is defined.
The execution loop, and what happens when things go wrong
Here is one run from prompt to result, with the places where it can retry, loop, pause, or stop.
This is the part I spent the most time on. A Golem run is a transparent loop, and each failure has a named place to land.
Self-correction
When a model produces bad output or bad tool arguments, the cheapest fix is often to tell it what was wrong and ask again. A decoder or tool does that by returning an error that wraps *model.ModelRetry:
if input.N <= 0 {
return "", &model.ModelRetry{Err: fmt.Errorf("n must be positive, got %d", input.N)}
}
The run delivers that message back to the model as feedback. You set separate budgets with WithOutputRetries and WithToolRetries, so a model cannot argue with your validator forever. Without a budget, the same error fails the run at a named stage with its cause intact.
A failure the model should simply see, like "that file does not exist," is a different thing: return &tool.Failed{Reason} and the model gets the reason as a tool result without spending retry budget.
Retries and fallbacks
Transient provider errors (408, 429, 5xx, transport faults) are retried with WithMaxAttempts. Backoff starts at 500 ms, doubles, and caps at 30 seconds unless you supply your own, jitter included. Tool failures and decode failures are never retried by this mechanism, since they are not transport problems.
When retrying is not enough, model.NewFallback(primary, backup) tries models in order. One rule for streams: a fallback only happens before the first fragment is forwarded. After the caller has seen output, switching models would replay it.
Limits that stop a runaway loop
golem.UsageLimit bounds input, output, and total tokens, model requests, tool calls, and, if you give the agent a price table, dollars. Tokens come from the provider's reported usage, requests and tool calls are counted by the run itself, and the check happens after each model response. A crossed bound fails the run at the usage stage with a typed UsageLimitError naming which dimension tripped.
Pricing is yours to supply through model.Price. Golem does not ship a table of model prices, because those go stale faster than a library release cycle.
Evidence survives failure
When a run fails on turn six, you still want turns one through five. Golem attaches them to the error as RunError.Partial: the conversation through the last completed model turn, the usage reported so far, and the counts of requests and tool executions. You can pass Partial.Messages to RunWithHistory and resume instead of starting over. For an agent that spent real money getting that far, this matters.
Pausing for a human
Some tools should not run without sign-off: deleting files, moving money, sending email. A tool can defer by returning an error wrapping *tool.Deferred, with a kind of tool.DeferApproval or tool.DeferExternal. The run then pauses cleanly. The other tool calls in that batch execute, and Run returns successfully with Result.Pending listing what is waiting. No further model call happens.
You resume with RunWithDeferredResults, passing the paused conversation and a decision for each pending call. Approved calls re-execute, denied calls tell the model it was denied, and external results are inserted verbatim. Because a pause is just a returned value, you can store it, show it in a UI, and resume it hours later from a different process.
Talking to the outside world: MCP
The Model Context Protocol has become the usual way to expose tools to agents, so Golem includes a client. mcp.NewStdio starts a server as a subprocess, mcp.NewHTTP connects to a remote endpoint, and mcp.AsTools turns the server's tool list into ordinary tool.Tool values with the server's own schema.
A server error flagged as isError comes back as a correctable ModelRetry, so the model can see the explanation and try again under the same retry budget. Protocol and transport failures fail the run instead of being quietly corrected. The client is synchronous, runs no goroutines, and respects your context on every call.
The guide for it says one thing I agree with: do not reach for MCP when a typed Go tool is one function away. A subprocess speaking JSON-RPC is more moving parts than a function call.
Documents without Python
Agents are constantly asked to read PDFs and Office files, and the usual answer is to shell out to a Python toolchain. Golem includes pure Go extractors instead:
pdfextractproduces Markdown from PDFs with reading order, multi-column layout, tables, and embedded images. Scanned pages can go to a vision model or Mistral OCR, with cost tracked per page.docextracthandles Word, Excel, PowerPoint, Markdown, CSV, and TSV, with an outline mode that returns a table of contents instead of the whole body so you do not spend your context window on a 200-page file.
Both are built on the standard library, which keeps the "zero dependencies" promise intact.
Testing agents without a model
Agent tests are often either flaky (real provider) or meaningless (mocked HTTP). Golem ships testmodel, a scripted in-memory model that plays back responses you queue and records the exact request the agent sent:
client := testmodel.New().Respond(
model.Response{Message: model.Message{Role: model.RoleAssistant, Content: "pong"}},
)
agent, _ := golem.New[struct{}, string](client, decoder)
result, err := agent.Run(context.Background(), golem.RunContext[struct{}]{}, "ping")
You can assert on three things: the normalized request, the ordered evidence in Result.Messages, and which stage failed. No network, no keys, no sleeps. The repository runs its suite with the race detector in CI, plus a fuzz target on the message JSON format so that anything Golem writes can be read back and rewritten byte for byte.
How it holds up in a real app
Herbie, a self-hosted chat and workflow app I built, runs its Go backend on Golem, and it is the main place I exercise the framework against real use.
One boundary deserves a mention: a browser that resumes a conversation sends the history back to you, and that history can say anything. It can include a forged system message or an image URL with a file:// scheme. golem.SanitizeHistory strips client-supplied system messages, drops URL parts that are not HTTP or HTTPS, repairs unresolved tool calls, and returns a report of what it removed, so the boundary can see what the client tried.
Where it stands
Golem is young. The repository started in August 2026 and the core execution contract is frozen while the surface around it grows. Everything I describe above is in the current release, and the public API changes only by addition on the way to v1.0.
If you are looking for a Go AI agent framework, here is how I would decide:
- Pick Golem if you want agents inside a Go service, typed dependencies, a dependency-free module, and tests that run offline.
- Pick something else if your agent is mostly prompt experimentation or you need a large ecosystem of prebuilt integrations today. Python is still ahead there, and I would not pretend otherwise.
The documentation site has a guide for each feature, and the examples/ directory has runnable programs for nearly all of them. Several run offline with no API key.
- Repository: github.com/abubakarsiddik31/golem
- Docs: abubakarsiddik31.github.io/golem
- API reference: pkg.go.dev/github.com/abubakarsiddik31/golem
Issues and pull requests are welcome. If you try it and something feels un-Go-like, I want to hear about it.
Frequently asked questions
What is Golem? Golem is an open-source AI agent framework for Go. It provides typed agents, tools with dependency injection, structured output, MCP support, and production controls like retries, usage limits, and human approval.
Does Golem have external dependencies? No. It is built on the Go standard library only.
Which LLM providers does it support? OpenAI, Anthropic, Google Gemini, AWS Bedrock, Azure OpenAI, and local models through Ollama, LM Studio, or any OpenAI-compatible endpoint.
Can I test agents without calling an API?
Yes. The testmodel package gives you a deterministic fake model, and several examples run fully offline.
Does it support MCP?
Yes. The mcp package is a client for stdio and streamable HTTP servers, and it exposes their tools as normal Golem tools.
About the author
Abu Bakar Siddik
Co-founder & Lead AI Engineer. He builds production LLM systems — agent orchestration, tool reliability, and private deployments for regulated environments — and takes on a limited number of consulting engagements at a time.
Follow along
New essays go up here first. Follow via RSS.
Related