Skip to content

ADR-0031: The chatbot routes to a Source Span; a cloud LLM reads the question, never the SDS

Status: Accepted — supersedes nothing. ADR-0002 and ADR-0012 stand in full. Date: 2026-09-08

Decisions

D-65 The Conversational surface routes to Safety-Critical Content and never composes it

A natural-language question resolves to a Source Span, displayed verbatim with its SDS Section, or to the chooser, or to Not Stated. The model decides where to point, never what the safety answer says. This is Routing.

D-66 Routing is two-tier: a client-side Routing Floor that is always present, and a cloud Query Intent tier that is not

The Routing Floor is lexical retrieval over the IndexedDB Corpus and runs with no network. The cloud tier normalises the question into Query Intent — colloquial Thai to product name, "ถุงมือ" to the PPE field — and its output is fed into the same Routing Floor index. With no signal, the Floor answers alone. The cloud tier is an upgrade, never a gate.

D-67 Below the Routing Threshold the answer is Not Stated, and the SDS Section may be named without asserting its content

Retrieval always has a top hit; that is not a reason to show one. Under the threshold the surface returns Not Stated and may point at the SDS Section that would hold such an answer, in Conversational voice, visibly distinct. It may not fill the Not Stated. The threshold is fixed at Curation, not tuned at runtime.

D-68 Routing never designates a Chemical Record

A chemical name in a question is handed to manual search, and D-46's chooser renders by product name, supplier and revision date. Routing selects nothing. Where the user reached the surface from a Chemical Record, Routing is scoped to that record and the scope is displayed.

D-69 Only Query Intent crosses the network; no Corpus text reaches the provider

The provider receives the user's question and returns structured intent. Source Spans, SDS Sections, Curated Translations and original documents never leave. Zero retention and no training on submitted text are required of the provider, and no history is placed in the prompt (D-09). The prototype therefore uses no provider embedding API, and retrieval is lexical on both tiers.

D-70 An Escalation Trigger fires client-side ahead of any network call, and Routing never selects a branch

A curated Thai and English phrase list runs in the browser before the cloud tier is contacted. On a match, Escalation and hotline 1669 render immediately and the user is handed to the surface. Routing delivers the user to the branch condition gate and never past it, pre-filling nothing — not spill volume, not exposure route — even where the question stated one.

D-71 Routing correctness is verified by a Routing Fixture of exact per-case outcomes

A curated set of Thai and English questions, each with one expected outcome: a named span, the chooser, or Not Stated. Every case must pass, and every case must also pass with the cloud tier dead. No case states a percentage or a threshold (D-30). Negative cases are mandatory: a chemical outside the Corpus, a question the SDS never answers, and a chemical name held by two supplier records.

Context

R-17 names an AI Chatbot with automatic language processing among the four technologies, R-01 frames the whole product as an AI assistant, and R-32 records that SDS data is not only scattered but hard to understand. ADR-0002 answered the generation half of that promise and answered it hard. It did not specify the retrieval half, and the poster's chatbot has sat in the spec since as one paragraph of permission (PRD §9) with no mechanism behind it.

The owner asked whether a cloud LLM with RAG could deliver that chatbot. Half of the answer was already yes and already written: D-17 permits generated Conversational Content and D-40 already places an online-only chatbot service on the server. The other half is the option ADR-0002 rejected by name.

This ADR exists because the difference between those halves is one design decision — whether the model writes the answer or points at it — and that decision was, until now, nowhere recorded.

The two-tier shape is forced by the problem statement rather than chosen for elegance. The moment of peak need is a spill, indoors, urgent, "frequently with no network signal". A chatbot that is the natural entry point and also the only surface that dies without a connection would concentrate failure precisely where D-06 exists to prevent it.

Decision

The Conversational surface becomes a router. It resolves questions to Source Spans that already exist in the Corpus, selecting nothing and writing no safety prose. A cloud LLM reads the user's question and returns structured intent; it never reads the SDS. Retrieval runs client-side over IndexedDB whether or not the cloud tier is reachable, so the surface degrades rather than dies. Every judgement the design needs — the confidence threshold, the emergency phrase list, the correctness fixture — is fixed once at Curation and read by a runtime that decides nothing.

Rejected options

These are the branches the owner was offered and declined, recorded so that the next agent does not re-propose them as improvements.

  • Generative answers grounded on retrieved spans — rejected. This is ADR-0002's own rejected option under a newer name, and its ruling stands verbatim: "Grounding reduces paraphrase drift; it does not eliminate it, and the residual failure lands on a person's hands in a spill." The worked case is in that document: for acetone, a model that swaps nitrile for latex "sounds identical and is wrong in a way that reaches skin."
  • Generated prose above the verbatim span — rejected. ADR-0012 already ruled on this exact layout: "the summary is what people would actually read, so a paraphrase error still reaches the user with the verbatim text serving as decoration."
  • Server-only routing, chatbot unavailable offline — rejected. It is the current §9 behaviour and it was acceptable while the surface only defined vocabulary. Once the surface is a route into First Aid, an online-only route is a safety path with a network dependency in it, and the problem statement says the signal is "frequently" absent at that moment.
  • Client-only routing, no cloud tier — rejected. Safe, and it surrenders the thing R-17 promised: ADR-0002 already named this shape, "reduces R-17's AI Chatbot to a synonym-aware search box."
  • Ranking the best-matching record and answering from it — rejected. It designates a canonical record, which D-45 forbids, and the failure is worse in a router than in a generator: the wrong supplier's span arrives verbatim, cited and Section-numbered, wearing full extractive authority. The repo's own acetone case is the proof — one sheet recommends nitrile gloves, another forbids storing with nitrile rubber.
  • Sending candidate spans to the provider for reranking — rejected. Better ranking, and it egresses The Plant's own supplier documents to a third party. Nobody on this effort has the authority to license them, and ADR-0001 rules that what only a company can supply is stubbed and named, not assumed.
  • Embedding the Corpus through a provider embedding API — rejected for the same reason at larger scale: it egresses the whole Corpus once rather than per query.
  • Cloud classification of urgency — rejected. It places a network round-trip in front of hotline 1669, which breaks D-06 and D-08 together at the one moment both exist for.
  • LLM-as-judge evaluation of routing quality — rejected. It measures against a model's opinion rather than the real baseline ADR-0017 requires, and it readmits a model to the safety loop through the test suite.

This ruling may not be re-decided

If a change contradicts this ADR: stop and raise it. Do not implement over it.

Explicit abuse warning. "We added RAG" is exactly the phrase that will later be used to reopen ADR-0002. Retrieval is not generation, and permission for the first is not permission for the second. Any build in which the Conversational surface emits a sentence of hazard, PPE, storage, spill or first-aid content that is not a verbatim Source Span has re-decided ADR-0002, whatever it is called, and this ADR is not the licence for it.

Specifically: do not let the cloud tier see a Source Span because ranking would improve; do not lower the Routing Threshold because a demo returned Not Stated; do not pre-fill a spill branch because the user "already said" the volume; do not make the Escalation Trigger a model call because a phrase list looks unsophisticated.

Consequences

What becomes easy. The poster's chatbot promise is deliverable without touching ADR-0002, and the licensing question about The Plant's documents never has to be asked. The cloud tier is cheap: short questions in, small structured intent out.

What becomes hard. Routing quality is now bounded by lexical retrieval over a small Corpus, because D-69 rules out provider embeddings. Expect Not Stated more often than a generative chatbot would produce, which is intended behaviour and is itself a finding for study objective R-04.2.

What is added to Curation. Three more artifacts, each a human judgement made once under review: the Routing Threshold, the Escalation Trigger phrase list, and the Routing Fixture. This is consistent with the structural rule of §5 — judgement happens at Curation and the runtime is deliberately dumb — but it is real cost on top of the half-day-to-a-day per document of D-58, and it is not per-document work.

What is not decided here. The provider and model, and the spend cap. No model is named anywhere in this repo and none is named here; it is recorded as a Declared Gap for the owner.

Coverage

UpstreamLanded inEvidenceNote
R-01D-65the "AI assistant" framing satisfied by routing rather than generationADR-0002 governs the generation half
R-05D-65, D-68the five capabilities are reached by Routing, not re-answered by it
R-17D-65, D-66, D-69"AI Chatbot with automatic language processing" resolved to a router with a cloud query tierADR-0002 weighed the same finding for generation