A decision surface, not a chatbot; designed at the infancy of LLM tooling. OpenAI GPT‑3.5 and Google PaLM 2 behind a federated retrieval layer; every response renders as four typed, independently-citable blocks — fact, inference, risk, action — with a human accountable at every beat. Scroll to follow a single Galliprant query from plain-language ask to customer-ready artifact.
What's the Galliprant dosing for senior dogs, and what should I confirm before recommending it to a clinic?
A query is the start of a workflow, not the end of one. The map below shows what happens between a salesperson asking and a customer receiving a structured answer — each lane holds a specific commitment.
This was the infancy of LLM products, with no established interaction patterns to lean on. I set the surface's design language by composing ElancoGPT from Uplook components so it read as native enterprise software from day one. Two carried most of the surface: the Alert, run in a persistent variant to hold the accuracy warning above the thread, and the Text Field with paired Button, restyled into the composer. The corpus was scoped deliberately: retrieval was tuned to the live needs of multiple key functional teams across Elanco (commercial, vet affairs, IT) rather than a generic knowledge base. At the infancy of LLMs, custom prompts were built around roles and personas, each one steering the model to render the information relevant to that user.
bg-amber-50border-l-amber-600text-smPinned above the thread instead of dismissible, so the accuracy contract never scrolls away.
bg-whiterounded-mdbg-elanco-600The stock pairing, restyled into a persistent composer; no new input component to maintain.
A field rep working a senior-heavy clinic types: "What's the Galliprant dosing for senior dogs, and what should I confirm before recommending it?" No prompt engineering, no keyword formatting — the interface meets them where they already think.
Fact, inference, risk, action: each block independently citable, copyable, and attachable to a downstream artifact. Rendered in EDS components, so the AI reads like the rest of the enterprise.
Every block ships with provenance: the FDA label on the fact, a confidence score on the inference, dated clinical guidance on the risk. A persistent warning names the model; every user is accountable for verifying each response before acting on it — ElancoGPT is an internal tool, so the human in the loop is an Elanco employee, not a customer. The Experimental surface names PaLM‑2 in the same chrome — federation stays invisible.
The action block copies straight into an email draft, with the dosing card and LFT reminder ready to send. Copy messages exports the transcript; Customise wraps a prompt into a shareable persona deeplink, so teams share links, not prompt lore.
Every block has provenance. Every block is independently exportable. The user controls which blocks travel to the customer.
One high-fidelity mockup, two live model demos, the shipped surfaces, and the responsive state boards; every screen built on EDS components.
The UX target the build was measured against: Elanco-branded, on-system, accountable by default.
The same surface, federated across OpenAI GPT‑3.5 and Google PaLM 2; the model swap is invisible to the user.
The desktop screens a session actually opens on: default, experimental, and the pages that let a user shape their own. Each shown at full width, on the same EDS chrome.
One system across every breakpoint: the full state walk from beta alert to streamed answer.
A chatbot that returns paragraphs is the lowest-value AI surface in an enterprise context. Fact / inference / risk / action made the AI a decision surface, not a reading surface.
Vet-affairs audit ran on a spreadsheet for the first six months. A purpose-built queue could have surfaced corpus gaps faster.