MCP servers  /  pai-bok

pai-bok

The Production AI Body of Knowledge, available to a coding agent over MCP. Answers come from authored reference material, not from the model's own recall.

Endpoint

https://bok.aisystemslab.org

Transport

HTTP MCP · JSON-RPC

Protocol version

2025-06-18

Auth

Bearer token, one per person

Example

How you would use it

You are building an agent. It calls a model, and sometimes the call hangs. You do not know how long to wait before giving up, or whether retrying will make things worse. So you ask, in the editor you are already working in.

You ask

How long should my agent wait for a model call, and should it retry?

Your agent calls recommend. This comes back

The decision

Timeout, retry and backoff policy for model and tool calls D15.M04

How long the loop waits, whether it tries again, and how it spaces attempts. The most-exercised reliability surface in an agent; every turn passes through it.

Four options, each with the failure it accepts

Capped exponential backoff with jitter T01

Spaces retries so that clients which failed together do not come back together. Coverage is partial: jitter appears nowhere else in the corpus, and the sources are one author at one vendor.

Client-side retry throttling by token bucket T02

A ceiling on total retry volume. A per-request limit of three still permits every request to retry three times at once; a budget does not.

Idempotency tokens for non-idempotent tool calls T03

Makes a retried side-effecting call safe. Without it, the other three are dangerous rather than merely imperfect.

Latency hedging by speculative parallel request T04

A second call to a faster model once the first passes a latency percentile, taking whichever returns first. Buys tail latency with inference cost.

First move

Write the timeout budget table: how long each hop is allowed to take, nested so that inner timeouts fit inside outer ones.

What it will not do

Tell you which one to pick. The corpus records what each option costs and runs no benchmark, so naming a winner would be inventing a result. You choose the failure you can live with.

If that is about to shape production code, the follow-up is cite D15.M04.T01. It returns four sources, marks which were read in full and which were only seen second-hand, and warns that the primary ones are all by a single author at a single vendor.

Every entry in the corpus answers in this shape: a field that already solved a reliability problem, what its solution becomes when the system runs on a model, and what the corpus still does not know.

Structure

Four levels, from broad to concrete. The same retry policy appears at each one.

L1

Discipline

The field the answer came from.

Reliability and resilience engineering

L3

Method

A commitment. Selecting one rules others out, which is what separates a method from a suggestion.

Timeout, retry and backoff policy

L4

Technology

The mechanism that carries the method out.

Capped exponential backoff with jitter

L5

Artifact

The thing the method produces.

Timeout budget table

There is no L2. The ten assurance pillars are tags that cut across the tree rather than a level within it; a single pillar applies to disciplines, methods, and technologies alike. Read as a tier between L1 and L3, consistent results will appear inconsistent.

The nine tools

The tools are in two groups. Building tools answer what to do next; curating tools answer what the corpus contains. Asking the wrong group returns a correct answer to a different question.

Building questions take the form “what should I do about X.” Curating questions take the form “what do we have on X.” Each example below is a real call against the corpus.

Building — you want a decision

recommend

The decision in front of you, what ships by default if it is skipped, every option with the failure it accepts, the first artifact to write, and what to monitor in production. It does not name a winner; the reason is below.

Example

“My agent calls a model and sometimes it hangs. How long do I wait, and do I retry?”

Returns the decision as D15.M04 with four options, each carrying the failure it accepts: capped backoff with jitter, retry throttling by token bucket, idempotency tokens, and latency hedging. First move: write the timeout budget table. No option is marked best.

compare_options

Every option under one method, side by side, with the recorded reasoning for each.

Example

compare_options D15.M04

Lays those four mechanisms alongside each other with the reasoning recorded for each, including the one offered as a replacement for a control another part of the corpus mandates. The conflict is shown rather than resolved.

Curating — you are asking about the corpus

find_practice

What the corpus holds on a problem and how well it holds it, including the thinness of coverage.

Example

“How well do we cover retry policy?”

Returns D15.M04 rated partial, with the reasoning: bounded retries, per-call timeouts and exponential backoff appear across four corpus principles, while jitter, retry budgets and percentile-derived timeouts return zero hits.

whats_covered

Whether a topic is in the corpus at all. Run it before adding anything.

Example

“prompt injection defence”

Verdict: covered. Nearest entries are context injection defence, an injection pattern library, and memory poisoning defence. It searches by meaning and by exact term together, because the two miss different things.

cite

Where a claim came from, whether it rests on a single source, and what has been corrected or withdrawn.

Example

cite D15.M04.T01

Four sources, each marked for whether it was read in full or only seen second-hand, plus an explicit warning that the primary sources are single-author and single-vendor. That warning is the reason the tool exists.

Escape hatches — when neither group fits

search

Raw retrieval in three modes: by meaning, by keyword across the full text, or by metadata filters alone.

Example

“reliability practice borrowed from an established engineering discipline”

A ranked list of addresses with level and coverage, from chaos engineering to dependency reliability budgeting. Pointers only, not content. Filters apply before ranking, so narrowing a search cannot discard the best match.

get_node

One entry in full, with what sits above and below it, and the raw source on request.

Example

get_node D15.M04

The entry in full, its parent discipline, and the five things beneath it: four mechanisms and one artifact, each with its own address.

Reporting back

report_gap

The corpus returned nothing on a topic that belongs in it. Writes to a queue a person reads.

Example

“Nothing here on retry storms amplified across nested agent layers.”

Written to the queue as a single record, stamped with the name the key was issued to, so the gap can be followed up with the person who hit it.

report_feedback

A recommendation was applied in a real system, with the outcome.

Example

“Used the token bucket ceiling. It held through a provider outage that would otherwise have amplified.”

Lands in the same queue. Reports describing a failure are read first, because a recommendation that broke in the field is the most useful thing in it.

Design constraints

No text generation on the server

Nothing inside the server generates text. One model call exists and it is an embedding, Titan Text Embeddings v2, which converts a question into 1024 numbers so that vector search can run. The function's permissions cover that model alone; it cannot invoke a text generator.

Everything returned is authored text or formatting applied to it. An embedding is still a model, so the claim is no text generation on the server, not the absence of AI. The client's own model does the talking.

No ranking of options

The corpus records the failure each option accepts. It runs no benchmarks.

Ranking them would introduce a finding the corpus never made and attach its authority to it. recommend states the decision and the cost of each branch; the choice stays with the person who knows the system.

Reports cannot alter the corpus

report_gap and report_feedback write to an inbox that no entry cites. A report is grounds to investigate, nothing more.

Amending the corpus requires a field note: applied work with verifiable evidence. A report sent over MCP does not meet that standard, and the architecture prevents it from becoming one.

Architecture

Serverless end to end. Nothing runs while nobody is asking: no container, no database process, no instance to patch. The whole service is a function, an object store and a vector index.

request ──> API Gateway ──> Lambda ──┬──> S3 Vectors      semantic search
                                     │
                                     └──> bok.db in /tmp  graph, fields, FTS5

Serverless vector search

A question is embedded once by Bedrock, then matched against S3 Vectors, which returns identifiers rather than content. There is no vector database to run, size or scale. Metadata filters apply before ranking, so narrowing a search cannot discard the best match.

SQLite as the graph

Everything else is one SQLite file of 10.6 MB holding every field, the taxonomy and a full-text index. The function pulls it from object storage once per container and opens it read only. Each entry stores its parent, so walking up to ancestors or down to children is an ordinary query rather than a graph engine.

What that buys

A warm keyword query returns in about 225 ms and touches no network at all, because the database is already on local disk. A semantic query takes about 460 ms, the difference being the single embedding call. Cold start is about 813 ms, paid once per container.

Permissions are the guarantee

The function's IAM policy grants exactly four things: read on one corpus snapshot, write on the inbox, query on one vector index, and invoke on one embedding model. Nothing else in the account is reachable from it.

That policy is what makes the constraints above structural rather than a matter of discipline. The server cannot generate text because it holds no permission to call anything that can.

Versioned snapshots

The corpus is served as an immutable dated snapshot, published only by a tagged release. An answer therefore belongs to a known version of the body of knowledge rather than to whatever it happened to say that day.

A container keeps the snapshot it downloaded until it expires, so a new release reaches the server at the next cold start rather than mid-conversation.

API Gateway rather than a bare function endpoint, because throttling, custom domains, WAF and access logging exist there and do not exist on one.

Reading the results

Provenance

Some entries rest on a single source. One records a claim withdrawn as incorrect. That history is retained rather than removed, on the view that a corpus concealing its corrections is harder to trust, not easier.

cite returns it, including whether each source was read in full or only seen second-hand. Check provenance before an entry shapes production work.

focus is not a quality score

Focus is proximity × gap multiplier: how tightly an entry couples to a running system, multiplied by how thinly it is covered.

A high score therefore indicates under-served and consequential. It does not indicate quality, importance, or endorsement. It shows where the corpus needs work, not what to adopt. Treating it as a ranking is the most common misreading.

Connecting

With a key issued, add the server to Claude Code:

claude mcp add --transport http pai-bok --scope user https://bok.aisystemslab.org --header "Authorization: Bearer YOUR_TOKEN"

Keep the word Bearer. It is easily dropped when substituting a token, and the server correctly rejects a bare token with a 401.

If a 401 appears, confirm Bearer still precedes the token before investigating further. This has already cost one person a debugging cycle.

Logging

Searches that return nothing are logged, with the question and the name of the requester, and expire after 180 days. This is deliberate. A question the corpus cannot answer is the most useful signal it produces, and it needs to reach the Lab rather than disappear. It is retention, so it is stated here rather than done quietly.

Limits

Rate

Two requests per second sustained, burst five. The ceiling sits on the endpoint rather than on each key, so a busy caller can consume it and others receive 429s. The accepted trade is that a per-key quota would mean a second authentication scheme beside the bearer token, which is more machinery than this size of service justifies.

Keys

One per person, issued and revoked individually, and compared in constant time on every request. Revocation takes effect within about a minute. The token is the identity: there is no login and no user record, which is weak as authentication and adequate as attribution, which is all it is used for.

Requesting access

There is no self-serve path. Keys are issued to fellows and collaborators, and each key belongs to a named person, so the Lab can see how the corpus is used and where it falls short.

Email [email protected] with the subject “MCP access — pai-bok”, stating who you are and what you are building. Other contact routes are on the Contact page.

The Lab does not publish tokens and does not reissue another person's.

Request access

State who you are and what you are building.

Email the Lab