How the factory belt works
A request (or a city need) arrives. The factory decides whether it can build it - cheaply by keyword-matching the need to its building-block machines, or smartly by asking the LLM autopilot. If yes, it builds a pipeline from those machines and runs it.
On the belt, each station does one job and passes a typed payload (a <PipelinePayload> of <Slice> rows) to the next - a station can read any earlier station's slice by name. Output from an untrusted station (the LLM, web search, external services) is checked at a trust boundary - validated against an XSD or a ContentType - before it flows on; trusted deterministic transforms just flow. The finished product carries a manifest slip (what it is + the route it took).
flowchart TD
Req[Request or city need] --> Decide{Build it?}
Decide -->|cheap keyword match| Build[Build pipeline]
Decide -->|smart LLM autopilot| Build
Build --> S1[Station 1]
S1 -->|typed payload| S2[Station 2]
S2 -->|typed payload| S3[Station 3]
S3 --> Out[Product plus manifest slip]
S1 -.->|untrusted output| Check{Trust boundary}
Check -->|untrusted - validate vs XSD or ContentType| Reject[Reject if bad]
Check -.->|trusted - just flow on| S2
flowchart TD Observe[Observe - Tool.CityStatus] --> Decide[Decide - what to build next] Decide --> Act[Act - build it or escalate a gap] Act --> City[(The living city)] City --> Observe
flowchart TD Eval["Eval part (XML) - trigger, scenario, tier-tagged assertions"] --> Det["Deterministic tier - stub model, real served stack"] Eval --> Qual[Quality tier - the real LLM offline] Det -->|"grades the PIPELINE - parses, validates, lands"| Grade[AssertionEvaluator - one vocabulary] Qual -->|grades the real ANSWER| Grade Grade -->|text assertions| Raw[Raw emission] Grade -->|XML assertions| Row[Landed row] Grade -->|lands-in-list| Store[(Storage re-read)] Grade --> Verdict["EvalVerdict - Pass, Fail or Degraded"] Verdict --> Result[(EvalResult row - queryable)]