AI Integration Layer¶
GTB provides the foundational chat package for interacting with LLMs.
Design Philosophy¶
pkg/chat is a thin, purpose-built abstraction for CLI tooling — not a general-purpose AI framework. Before selecting this approach, a comprehensive evaluation of alternatives (LangChain Go, go-openai, vercel/ai-sdk, and ~10 others) was performed.
The conclusion was clear: no existing library matched GTB's specific requirements:
- Minimal interface surface — the
ChatClientinterface exposes exactly four methods:Add,Chat,Ask, andSetTools. Downstream code never needs to know which provider is active. - Tool calling + structured output — both capabilities are required together across providers, a combination that most thin wrappers do not handle uniformly.
- CLI-first design — features like token chunking, history management, and subprocess providers (e.g.,
ProviderClaudeLocal) are CLI concerns that general AI frameworks don't address. - Extensible without forking — the
RegisterProviderregistry allows third-party packages to add providers without modifyingpkg/chat. - Testable by default — generated mocks in
pkg/mocks/chatallow downstream applications to stub the entire AI layer without network calls.
This positioning makes pkg/chat a "right-sized" component: large enough to solve real provider-abstraction complexity, small enough that its full interface fits on a single screen.
Core Abstractions¶
chat.ChatClient¶
The ChatClient interface is the primary entry point. It abstracts the differences between all supported providers:
type ChatClient interface {
Add(ctx context.Context, prompt string, media ...Media) error
Ask(ctx context.Context, question string, target any, media ...Media) error
Chat(ctx context.Context, prompt string, media ...Media) (string, error)
SetTools(tools []Tool) error
Usage() Usage
}
All five built-in providers implement this interface:
- OpenAI (and compatible APIs via
ProviderOpenAICompatible) - Google Gemini
- Anthropic Claude (API)
- Claude Local — via the locally installed
claudeCLI binary
Structured Output¶
The chat package supports structured output via JSON schemas. When adding new providers or improving existing ones, ensure that schema validation remains consistent across all backends. The GenerateSchema[T]() helper generates a *jsonschema.Schema from any Go type.
Multimodal Input¶
Add, Ask, Chat and StreamChat take a trailing variadic of Media — images (and, on Gemini, PDF and A/V) sent alongside the text prompt. The parameter is variadic, so text-only calls are unchanged and the media path never affects them.
img, _ := os.ReadFile("frame.jpg")
// Structured output over an image.
var out struct{ Description string `json:"description"` }
err := client.Ask(ctx,
"Describe this photo as JSON {\"description\": string}", &out,
chat.Media{Data: img}) // MIMEType optional — sniffed from the bytes
// Or free-form, with more than one attachment.
answer, err := client.Chat(ctx, "Which of these is better lit?",
chat.Media{Data: imgA}, chat.Media{Data: imgB})
Attachments are placed after the prompt text in the same user turn, in the order given.
The Media type¶
type Media struct {
MIMEType string // OPTIONAL — see below
Data []byte // the raw attachment bytes
}
MIMEType is an optional cross-check, not an override. The type is always sniffed from Data, and the sniffed type is what gets sent. If you set MIMEType, it must match the sniffed family (e.g. image/*) or the attachment is rejected — this catches disguised content (a ZIP labelled image/png) and your own bugs. Leave it empty to let detection do the work.
Safety filtering¶
Every attachment is validated before any network call:
- Sniff the type from the bytes with
net/http.DetectContentType(never from a filename). - Reconcile against a declared
MIMEType(family mismatch → rejected). - Allowlist — the sniffed type must be a recognised media type; anything else (executables, scripts, archives, HTML,
application/octet-stream) is refused. - Provider support — the type must be accepted by the selected provider (matrix below).
- Limits — at most 16 attachments per request, 20 MiB each.
Nothing is uploaded on failure. Two error sentinels distinguish the causes (test with errors.Is):
| Error | Meaning |
|---|---|
ErrMediaRejected |
Failed the safety filter: empty, oversize, over-count, an unidentifiable/disallowed type, or a declared-vs-sniffed mismatch. |
ErrMediaUnsupported |
The type is valid but the selected provider (or model) doesn't accept it — including any media sent to ProviderClaudeLocal. |
Provider support matrix¶
| Provider | Images | Audio/Video | |
|---|---|---|---|
| Gemini | ✅ | ✅ | ✅ |
| Claude | ✅ | ✅ | — |
| OpenAI (+ compatible) | ✅ | ✅ | — |
| Claude Local | — | — | — |
Each type maps to the provider's native shape: images to a Gemini inline blob / Claude base64 image block / OpenAI image_url data URI, and PDF to a Gemini inline blob / Claude base64 document block / OpenAI file part. The v1 accepted range is what stdlib detection can positively identify — common images, PDF, and the common A/V containers (mp4/webm/avi; mp3/wav/ogg/aiff). Long-tail formats stdlib cannot name (mov, flv, wmv, flac, m4a) are rejected until a richer sniffer lands.
Persistence¶
Media survives snapshots. When a ConversationStore (e.g. FileStore) saves a snapshot, large media in the messages is externalised to a content-addressed cache (<store>/media/<sha256>, deduplicated, and encrypted with the store key when one is set) and replaced with a short reference; Load resolves the references back, so a restored conversation keeps its attachments and the snapshot file itself stays small. This is generic (provider-agnostic) and fully reversible — a provider's Restore sees byte-identical messages.
Known limitations (v1)¶
- Audio/video is Gemini-only. Claude and OpenAI take images and PDF.
- Long-tail A/V formats stdlib cannot sniff (
mov,flv,wmv,flac,m4a) are rejected until a richer sniffer is swapped in behind the detection choke point.
Developing for the AI Layer¶
Adding a New Provider¶
Providers self-register via init() using the RegisterProvider extension point. To add a new AI provider from any package:
- Implement the
ChatClientinterface. - Register the implementation with a unique
Providerconstant name:
// myprovider/provider.go
package myprovider
import (
"context"
"gitlab.com/phpboyscout/go-tool-base/pkg/chat"
"gitlab.com/phpboyscout/go-tool-base/pkg/props"
)
func init() {
chat.RegisterProvider("my-backend", newMyBackend)
}
func newMyBackend(ctx context.Context, p *props.Props, cfg chat.Config) (chat.ChatClient, error) {
return &MyBackendClient{token: cfg.Token, baseURL: cfg.BaseURL}, nil
}
After importing your package (e.g., via a blank import in main.go), chat.New(ctx, p, chat.Config{Provider: "my-backend"}) routes to your factory. No changes to pkg/chat are required.
Built-in providers follow the same pattern — each file registers itself in its own init():
| File | Registers |
|---|---|
openai.go |
ProviderOpenAI, ProviderOpenAICompatible |
claude.go |
ProviderClaude |
gemini.go |
ProviderGemini |
claude_local.go |
ProviderClaudeLocal |
Testing AI Logic¶
We use a "Golden File" approach for testing complex AI prompts and responses. This ensures that changes to our implementation don't unintentionally alter the prompt structure sent to the models.
Mocking in Consumer Apps¶
We provide generated mocks in pkg/mocks/chat to allow downstream applications to test their AI-powered features without making real network calls. The mock satisfies ChatClient and can be configured to return canned responses or return errors.
mockClient := mocks.NewChatClient(t)
mockClient.EXPECT().Chat(mock.Anything, "analyze this").Return("looks good", nil)