MCP servers / pai-bok
pai-bok
The Production AI Body of Knowledge, available to a coding agent over MCP. Answers come from authored reference material, not from the model's own recall.
Endpoint
https://bok.aisystemslab.org
Transport
HTTP MCP · JSON-RPC
Protocol version
2025-06-18
Auth
Bearer token, one per person
Example
How you would use it
You are building an agent. It calls a model, and sometimes the call hangs. You do not know how long to wait before giving up, or whether retrying will make things worse. So you ask, in the editor you are already working in.
You ask
How long should my agent wait for a model call, and should it retry?
Your agent calls recommend. This comes back
The decision
Timeout, retry and backoff policy for model and tool calls D15.M04
How long the loop waits, whether it tries again, and how it spaces attempts. The most-exercised reliability surface in an agent; every turn passes through it.
Four options, each with the failure it accepts
Capped exponential backoff with jitter T01
Spaces retries so that clients which failed together do not come back together. Coverage is partial: jitter appears nowhere else in the corpus, and the sources are one author at one vendor.
Client-side retry throttling by token bucket T02
A ceiling on total retry volume. A per-request limit of three still permits every request to retry three times at once; a budget does not.
Idempotency tokens for non-idempotent tool calls T03
Makes a retried side-effecting call safe. Without it, the other three are dangerous rather than merely imperfect.
Latency hedging by speculative parallel request T04
A second call to a faster model once the first passes a latency percentile, taking whichever returns first. Buys tail latency with inference cost.
First move
Write the timeout budget table: how long each hop is allowed to take, nested so that inner timeouts fit inside outer ones.
What it will not do
Tell you which one to pick. The corpus records what each option costs and runs no benchmark, so naming a winner would be inventing a result. You choose the failure you can live with.
If that is about to shape production code, the follow-up is cite D15.M04.T01. It returns four sources, marks which were read in full and which were only seen second-hand, and warns that the primary ones are all by a single author at a single vendor.
Every entry in the corpus answers in this shape: a field that already solved a reliability problem, what its solution becomes when the system runs on a model, and what the corpus still does not know.
Structure
Four levels, from broad to concrete. The same retry policy appears at each one.
L1
Discipline
The field the answer came from.
Reliability and resilience engineering
L3
Method
A commitment. Selecting one rules others out, which is what separates a method from a suggestion.
Timeout, retry and backoff policy
L4
Technology
The mechanism that carries the method out.
Capped exponential backoff with jitter
L5
Artifact
The thing the method produces.
Timeout budget table
There is no L2. The ten assurance pillars are tags that cut across the tree rather than a level within it; a single pillar applies to disciplines, methods, and technologies alike. Read as a tier between L1 and L3, consistent results will appear inconsistent.
The nine tools
The tools are in two groups. Building tools answer what to do next; curating tools answer what the corpus contains. Asking the wrong group returns a correct answer to a different question.
Building questions take the form “what should I do about X.” Curating questions take the form “what do we have on X.” Each example below is a real call against the corpus.
Building — you want a decision
recommend
The decision in front of you, what ships by default if it is skipped, every option with the failure it accepts, the first artifact to write, and what to monitor in production. It does not name a winner; the reason is below.
Example
“My agent calls a model and sometimes it hangs. How long do I wait, and do I retry?”
Returns the decision as D15.M04 with four options, each carrying the failure it accepts: capped backoff with jitter, retry throttling by token bucket, idempotency tokens, and latency hedging. First move: write the timeout budget table. No option is marked best.
compare_options
Every option under one method, side by side, with the recorded reasoning for each.
Example
compare_options D15.M04
Lays those four mechanisms alongside each other with the reasoning recorded for each, including the one offered as a replacement for a control another part of the corpus mandates. The conflict is shown rather than resolved.
Curating — you are asking about the corpus
find_practice
What the corpus holds on a problem and how well it holds it, including the thinness of coverage.
Example
“How well do we cover retry policy?”
Returns D15.M04 rated partial, with the reasoning: bounded retries, per-call timeouts and exponential backoff appear across four corpus principles, while jitter, retry budgets and percentile-derived timeouts return zero hits.
whats_covered
Whether a topic is in the corpus at all. Run it before adding anything.
Example
“prompt injection defence”
Verdict: covered. Nearest entries are context injection defence, an injection pattern library, and memory poisoning defence. It searches by meaning and by exact term together, because the two miss different things.
cite
Where a claim came from, whether it rests on a single source, and what has been corrected or withdrawn.
Example
cite D15.M04.T01
Four sources, each marked for whether it was read in full or only seen second-hand, plus an explicit warning that the primary sources are single-author and single-vendor. That warning is the reason the tool exists.
Escape hatches — when neither group fits
search
Raw retrieval in three modes: by meaning, by keyword across the full text, or by metadata filters alone.
Example
“reliability practice borrowed from an established engineering discipline”
A ranked list of addresses with level and coverage, from chaos engineering to dependency reliability budgeting. Pointers only, not content. Filters apply before ranking, so narrowing a search cannot discard the best match.
get_node
One entry in full, with what sits above and below it, and the raw source on request.
Example
get_node D15.M04
The entry in full, its parent discipline, and the five things beneath it: four mechanisms and one artifact, each with its own address.
Reporting back
report_gap
The corpus returned nothing on a topic that belongs in it. Writes to a queue a person reads.
Example
“Nothing here on retry storms amplified across nested agent layers.”
Written to the queue as a single record, stamped with the name the key was issued to, so the gap can be followed up with the person who hit it.
report_feedback
A recommendation was applied in a real system, with the outcome.
Example
“Used the token bucket ceiling. It held through a provider outage that would otherwise have amplified.”
Lands in the same queue. Reports describing a failure are read first, because a recommendation that broke in the field is the most useful thing in it.
Design constraints
No text generation on the server
Nothing inside the server generates text. One model call exists and it is an embedding, Titan Text Embeddings v2, which converts a question into 1024 numbers so that vector search can run. The function's permissions cover that model alone; it cannot invoke a text generator.
Everything returned is authored text or formatting applied to it. An embedding is still a model, so the claim is no text generation on the server, not the absence of AI. The client's own model does the talking.
No ranking of options
The corpus records the failure each option accepts. It runs no benchmarks.
Ranking them would introduce a finding the corpus never made and attach its authority to it. recommend states the decision and the cost of each branch; the choice stays with the person who knows the system.
Reports cannot alter the corpus
report_gap and report_feedback write to an inbox that no entry cites. A report is grounds to investigate, nothing more.
Amending the corpus requires a field note: applied work with verifiable evidence. A report sent over MCP does not meet that standard, and the architecture prevents it from becoming one.
Architecture
Serverless end to end. Nothing runs while nobody is asking: no container, no database process, no instance to patch. The whole service is a function, an object store and a vector index.
request ──> API Gateway ──> Lambda ──┬──> S3 Vectors semantic search
│
└──> bok.db in /tmp graph, fields, FTS5
Serverless vector search
A question is embedded once by Bedrock, then matched against S3 Vectors, which returns identifiers rather than content. There is no vector database to run, size or scale. Metadata filters apply before ranking, so narrowing a search cannot discard the best match.
SQLite as the graph
Everything else is one SQLite file of 10.6 MB holding every field, the taxonomy and a full-text index. The function pulls it from object storage once per container and opens it read only. Each entry stores its parent, so walking up to ancestors or down to children is an ordinary query rather than a graph engine.
What that buys
A warm keyword query returns in about 225 ms and touches no network at all, because the database is already on local disk. A semantic query takes about 460 ms, the difference being the single embedding call. Cold start is about 813 ms, paid once per container.
Permissions are the guarantee
The function's IAM policy grants exactly four things: read on one corpus snapshot, write on the inbox, query on one vector index, and invoke on one embedding model. Nothing else in the account is reachable from it.
That policy is what makes the constraints above structural rather than a matter of discipline. The server cannot generate text because it holds no permission to call anything that can.
Versioned snapshots
The corpus is served as an immutable dated snapshot, published only by a tagged release. An answer therefore belongs to a known version of the body of knowledge rather than to whatever it happened to say that day.
A container keeps the snapshot it downloaded until it expires, so a new release reaches the server at the next cold start rather than mid-conversation.
API Gateway rather than a bare function endpoint, because throttling, custom domains, WAF and access logging exist there and do not exist on one.
Reading the results
Provenance
Some entries rest on a single source. One records a claim withdrawn as incorrect. That history is retained rather than removed, on the view that a corpus concealing its corrections is harder to trust, not easier.
cite returns it, including whether each source was read in full or only seen second-hand. Check provenance before an entry shapes production work.
focus is not a quality score
Focus is proximity × gap multiplier: how tightly an entry couples to a running system, multiplied by how thinly it is covered.
A high score therefore indicates under-served and consequential. It does not indicate quality, importance, or endorsement. It shows where the corpus needs work, not what to adopt. Treating it as a ranking is the most common misreading.
Connecting
With a key issued, add the server to Claude Code:
claude mcp add --transport http pai-bok --scope user https://bok.aisystemslab.org --header "Authorization: Bearer YOUR_TOKEN"
Keep the word Bearer. It is easily dropped when substituting a token, and the server correctly rejects a bare token with a 401.
If a 401 appears, confirm Bearer still precedes the token before investigating further. This has already cost one person a debugging cycle.
Logging
Searches that return nothing are logged, with the question and the name of the requester, and expire after 180 days. This is deliberate. A question the corpus cannot answer is the most useful signal it produces, and it needs to reach the Lab rather than disappear. It is retention, so it is stated here rather than done quietly.
Limits
Rate
Two requests per second sustained, burst five. The ceiling sits on the endpoint rather than on each key, so a busy caller can consume it and others receive 429s. The accepted trade is that a per-key quota would mean a second authentication scheme beside the bearer token, which is more machinery than this size of service justifies.
Keys
One per person, issued and revoked individually, and compared in constant time on every request. Revocation takes effect within about a minute. The token is the identity: there is no login and no user record, which is weak as authentication and adequate as attribution, which is all it is used for.
Requesting access
There is no self-serve path. Keys are issued to fellows and collaborators, and each key belongs to a named person, so the Lab can see how the corpus is used and where it falls short.
Email [email protected] with the subject “MCP access — pai-bok”, stating who you are and what you are building. Other contact routes are on the Contact page.
The Lab does not publish tokens and does not reissue another person's.