Briton Rivière, Daniel in the Lions' Den, 1872
Briton Rivière, Daniel in the Lions' Den, 1872

How our AI agents handle 70 tickets a day

A small service desk team drowning in tickets, and the multi-agent system I designed so most of them never reach a human. Agents that answer, interview, write and classify, all sharing one state.

agentslanggraphawsarchitecture

Seventy tickets a day. On bad days, a hundred. The service desk team handling them was small.

The usual answers are more people or faster people. My team and I proposed a third one: an agent that sits in front of the queue. Once the idea was approved I took it over. I spent a week reading real tickets, then designed the architecture, picked the stack, and started building.

What was actually in the queue

Three kinds of tickets, and only one of them needed a human.

Service desk queue70 today
6 tickets
  • How do I get access to the reporting dashboard?in the docs
  • VPN disconnects every few minutesknown fix
  • Errorno context
  • Where is the latest rate sheet?in the docs
  • Export not workingno context
  • Password reset link expiredknown fix

The first two had answers sitting in Confluence or in a ticket closed three weeks earlier. An engineer’s whole contribution was finding them.

The third kind was the real cost. A ticket that named an error and nothing else: no steps, no process, no expected result. The first reply was never a fix, it was a question. Then waiting. Then another question. A four-minute problem stretched across two days of round trips.

So the problem wasn’t speed. The input was malformed. Two goals came out of that: answer what’s answerable before it reaches a human, and make sure everything else arrives complete.

Why a chatbot would make it worse

A RAG chatbot over the docs covers the first two kinds and fails the third, badly. It answers until it can’t, then apologizes and links the same empty form. The user already explained the problem once, to a bot, for nothing. What do they type into the form now? The error name. Nothing else.

So “I couldn’t help” can’t be a dead end. The failure path is where the system has to do its most valuable work.

Agents that share one state

The whole system is a single LangGraph graph running models on Amazon Bedrock. Each agent has one job, and they all read and write the same state, so none of them ever asks the user something another one already knows.

One LangGraph graphHover a step to see what it does
in parallelsearch againUser questionRAG agentAnswer in chatTicket agentDrafterClassifierShared stateHuman reviewJira ticket

The RAG agent answers from a Bedrock Knowledge Base fed by Confluence and by every resolved ticket. The docs say how things should work. The resolved tickets say what people actually get wrong, in the same words users type. Indexing both made retrieval noticeably better than either alone.

The ticket agent is the part I’d defend hardest. When there’s no answer, it doesn’t apologize. It offers to open a ticket and interviews the user, one question at a time. This is what that feels like from the user’s side. Try the buttons at the end.

Service desk assistant

The export is failing

RAG agentno match in docs or resolved tickets
Ticket agentasking only for what's missing

Not in the docs yet, so let's open a ticket. Which export, and what happened?

Monthly report, last quarter, Export CSV. It times out after 30 seconds

Drafter + Classifierin parallel, on the same state
DraftMonthly report export times out on long rangesReporting · P3
Human reviewthe graph is paused until you decide

Mention it only fails for ranges over a month

Updated the description. Ready to create it?

Cancelled. Nothing was written to Jira.

Jira toolticket created

Done. SD-2047 is open, here's the link.

Nobody writes a good bug report into an empty box. Almost anyone can answer a few direct questions.

Every agent has the same shape: a Bedrock model, a prompt, and only the tools its job needs. The RAG agent can call its search tool as many times as it wants before answering, which is the loop in the graph above. The drafter has no tools at all, just a typed output.

agents.tsts
1import { createAgent, tool } from "langchain";
2import { z } from "zod";
3
4const searchDocs = tool(
5 async ({ query }) => knowledgeBase.invoke(query),
6 {
7 name: "search_docs",
8 description: "Search Confluence and resolved tickets",
9 schema: z.object({ query: z.string() }),
10 },
11);
12
13// Calls search_docs as many times as it needs,
14// then answers or admits it can't.
15export const ragAgent = createAgent({
16 model: bedrock,
17 tools: [searchDocs],
18 systemPrompt: "Only answer from search_docs results.",
19});
20
21export const ticketAgent = createAgent({
22 model: bedrock,
23 tools: [askUser],
24 systemPrompt: "Ask for missing details, one at a time.",
25});
26
27// No tools: reads the conversation, returns a typed draft
28export const drafter = createAgent({
29 model: bedrock,
30 responseFormat: TicketDraft,
31 systemPrompt: "Write the ticket from the conversation.",
32});

Once nothing is missing, the agent drafts the ticket from what the user already said. Their job shrinks from write a report to confirm this is right.

Service Desk / SD-2045

Export not working

Waiting for customer
Description
Steps to reproduce

Not provided

Actual result

doesn't work

Expected result

Not provided

4 comments asking for details · open for 2 days
Service Desk / SD-2047

Monthly report export times out on long ranges

Ready for triage
Description
Steps to reproduce

Reports → Monthly → filter by last quarter → Export CSV

Actual result

“Request timed out” after about 30 seconds

Expected result

A CSV download, as with shorter date ranges

Drafted by the assistant, approved by the reporter · no follow-up questions

Once the interview is done, two agents start at the same time. The drafter writes the title and description. The classifier sets category and priority. Neither waits for the other: both read the same conversation from the shared state, each writes its own field, and the review step only runs when both are finished. The ticket is ready in the time the slower of the two takes, not the sum of both.

The classifier works with plain rules, not model judgment. Routing affects SLAs and on-call, so it has to be deterministic and auditable. When someone asks why a ticket landed somewhere, the answer is a rule, not a prompt.

Wiring it all together is where LangGraph earns its place. The shared state, every node, the parallel branches and the checkpointer live in one file you can read top to bottom.

graph.tsts
1// One state object, shared by every node
2const TicketState = Annotation.Root({
3 ...MessagesAnnotation.spec,
4 draft: Annotation<TicketDraft>,
5 labels: Annotation<TicketLabels>,
6});
7
8const checkpointer = PostgresSaver.fromConnString(DB_URL);
9
10export const graph = new StateGraph(TicketState)
11 .addNode("rag", ragAgent)
12 .addNode("interview", ticketAgent)
13 .addNode("draft", drafter)
14 .addNode("classify", classifier)
15 .addNode("review", humanReview)
16 .addEdge(START, "rag")
17 .addConditionalEdges("rag", answeredOrTicket)
18 // Fan out: both branches run at the same time
19 .addEdge("interview", "draft")
20 .addEdge("interview", "classify")
21 // Fan in: review waits for both to finish
22 .addEdge(["draft", "classify"], "review")
23 .addEdge("review", END)
24 .compile({ checkpointer });

Nothing is written without a human

Before anything reaches Jira, the user sees the draft and can ask for changes, cancel it, or create it. The agent will sometimes misread intent. What it can’t do is file a ticket while being wrong.

In LangGraph that’s an interrupt. The graph stops mid-run, saves its state, and resumes from that exact line when the user answers, ten seconds or eighteen hours later.

review.tsts
1export async function humanReview(state: TicketState) {
2 // Stops here. Resumes when the user clicks a button,
3 // ten seconds or eighteen hours later.
4 const decision = interrupt({ draft: state.draft });
5
6 if (decision.action === "cancel") return { draft: null };
7
8 const ticket = await jira.createIssue(decision.draft);
9 return { ticketUrl: ticket.url };
10}

If your agent can pause for a human, it needs a real database, and it isn’t the vector store. A paused graph has to survive cold starts and deploys, which is why the graph above compiles with a Postgres checkpointer.

The architecture

Everything runs on AWS, defined with SST and written in TypeScript.

One AWS accountNothing leaves it. Hover a layer.
Entry
API GatewayLambda
Agents
LangGraphLangChainBedrock models
Knowledge
Bedrock Knowledge BaseConfluenceNotion
State
Aurora PostgreSQLRedisDynamoDB
Observability
Langfuse on ECS

One constraint shaped all of it. The end client, NFTYDoor, is a US lending platform, so every prompt, trace and retrieved chunk can contain financial data. Managed observability was out. Langfuse runs self-hosted on ECS, inside the same account as everything else.

Compliance changes the question you ask. Not what’s the fastest way to wire this, but where does this data physically go, and who can read it. Slower to build. It’s also the difference between a demo and something a regulated company runs in production.

Every answer also takes a thumbs up or down with an optional comment. Without that signal, a drop in tickets could mean good answers or users giving up. The thumbs-down comments became the most useful data in the system: each one points at a specific gap in the knowledge base.

What changed

Ticket volume dropped, and that’s the least interesting part. The tickets that still get created arrive complete: steps, exact error, expected result, the right category and priority. The team stopped opening tickets just to ask what the user meant.

The goal was never fewer tickets. It was fewer broken ones, and giving a small team back the hours they spent on work that never needed them.

If you build something similar, decide two things before picking a model: where your agent’s state lives, and what it does the moment it doesn’t know. Most of the value in this project came from the second one.