Writing · Applied AI

Building an Agentic Support System for Complex Connected Products

How I structure support agents to investigate product evidence, use tools carefully and hand useful work to a person when the case needs one.

The easiest support-agent demonstration starts with a clean question and ends with a clean answer. Real connected products rarely offer either.

A customer reports that a product is “offline.” That description might point to power, installation, local radio conditions, firmware state, credentials, an upstream service or a mobile interface showing stale state. It might also describe expected behavior using the wrong word. Treat the first message as one observation from one point in a larger system, not as the diagnosis.

This distinction changes the architecture. Connected-product support is a partially observable investigation across customer intent, physical installation, device state, firmware, connectivity, cloud services and operating history. Retrieval can find relevant knowledge, but a document is only one input to deciding which evidence matters now. A useful support agent must form and revise hypotheses, choose the next question or observation, remain inside its authority and recognize when uncertainty requires a handoff.

Safe diagnostic progress is the objective. Autonomy is incidental.

Support is a state-estimation problem

Software support often begins with logs, reproducible state and a release that can be identified precisely. A connected product adds physical conditions and several independent clocks. The device may have one view of the world, the cloud another, and the user’s phone a third. Each observation can be accurate for the moment in which it was captured and still mislead the investigation now.

An agent therefore needs a small but explicit investigation record. It should distinguish what was reported from what was observed, and both from what has merely been inferred.

A minimal evidence model for connected-product support
Record typeExampleHow the system should treat it
Customer report“The device has been offline since yesterday.”Preserve the wording and time context; do not silently convert it into a technical fact.
Observed stateThe service last received a valid heartbeat at a recorded time.Attach provenance, timestamp and any known limits of the observation.
Derived stateThe device is probably unable to reach the service.Label it as an inference and retain competing explanations.
Action resultA bounded connectivity test returned a specific response.Record the tool, scope, result and whether the action changed state.

This structure sounds bureaucratic until a case becomes ambiguous. Then it prevents a plausible early story from hardening into “what happened.” It also creates a useful audit trail when the case moves to another person.

Build the system around four operating artifacts

The chat transcript is not the support system. A production design needs four durable artifacts that remain coherent while the conversation changes.

Four artifacts behind an investigative support agent
ArtifactWhat it carriesWhat it prevents
Evidence ledgerReports, observations, source, time and limitationsInference hardening into fact.
Hypothesis setActive explanations, contradictions and discriminating next stepsPremature convergence on the first plausible story.
Permission envelopeCase scope, authorized tools, action class and approval stateDiagnostic progress drifting into unauthorized control.
Escalation packetCurrent case state, ruled-out paths, changes made and remaining riskA human receiving a transcript instead of an investigation.

These artifacts create separation of concerns. The model may propose a hypothesis, but it does not redefine the evidence. It may request a tool, but the permission envelope decides whether the request is valid. It may draft a handoff, but the packet is assembled from the record rather than from conversational memory.

This also makes the system replaceable at the right layer. A better model can improve judgment without changing authorization. A better diagnostic tool can improve evidence without rewriting the interaction. Governance becomes part of the architecture rather than a prompt instruction.

The investigation loop

An effective agent can be organized around a repeatable loop instead of a long prompt. Research such as ReAct studies the interleaving of reasoning and actions that gather information from an environment. In support, that pattern needs an operating layer around it: evidence provenance, explicit authorization, recovery and designed escalation.

1. Establish the intended outcome

“Make it work” is not always specific enough. Is the user trying to complete an installation, restore a previously working product, understand a notification, or confirm that a safety function is available? The same observed state can require a different next step depending on the intended outcome.

2. Build the smallest useful state picture

Gather the minimum evidence that can separate the leading hypotheses. This may include the product identity, relevant version, recent state, environment and the timing of the symptom. Do not collect every available field simply because a tool exposes it.

3. Keep competing hypotheses alive

The agent should resist premature convergence. Each active hypothesis needs supporting evidence, contradictory evidence and a next discriminating observation. A hypothesis with no possible falsifying observation is not yet useful.

4. Choose the next step by information value

The best next question is the one most likely to change the investigation. A screenshot can help if the wording distinguishes two states. A generic setup questionnaire adds nothing when the service has already observed a successful provisioning event.

5. Update confidence and risk

New evidence changes both likelihood and consequence. A low-confidence diagnosis can still justify a reversible observation; it should not justify a high-impact action. Confidence in the diagnosis and authority to act are separate variables.

6. Resolve, continue or escalate

The loop ends when the evidence supports a safe resolution, when another observation is worth collecting, or when a human is better placed to continue. Escalation should be a designed outcome, not an unstructured fallback.

The goal is to reduce uncertainty without hiding what remains unknown.

Ask fewer, better questions

Many support flows externalize the investigation burden. They ask the customer for model, version, installation date, network type, screenshots, serial number and a detailed symptom history before using any information already available to the organization.

An agent should instead rank questions by three properties:

  1. Discrimination: Will the answer meaningfully separate the active hypotheses?
  2. Answerability: Can the customer observe and report it reliably?
  3. Burden: Is the effort proportionate to the information gained?

Consider a hypothetical device that stopped reporting after a configuration change. Asking the customer to “describe the network” is broad and difficult to grade. Confirming whether a specific indicator changed immediately after the configuration step may divide the problem between local state and upstream connectivity with much less effort.

The agent should also explain why a question matters when the burden is not obvious. A short reason builds trust and helps the user supply the right observation without teaching them an internal troubleshooting tree.

Authority belongs outside the model

Giving an agent tools changes the product from an answer generator into an actor. The design unit is no longer only the model response; it is the combination of model, tool interface, authorization, observation and recovery path.

Tool classes require different controls
Tool classTypical purposeRequired discipline
Read-only observationInspect current state or historyMinimize data, record provenance, enforce tenant and case scope.
Bounded diagnosticRun a safe testDeclare scope, rate limits, expected effects and interpretation limits.
Reversible changeRefresh or reapply a known-safe settingRequire authorization, log before/after state and provide rollback.
High-impact actionChange behavior with safety, security or service consequencesKeep behind explicit human approval or outside the agent boundary.

Tool output is evidence, not truth. A service can return stale data. A “successful” command can mean only that a request was accepted. A missing record can be caused by scope, retention, synchronization or permissions. The tool contract should explain these limits in terms the agent can use.

The model should never infer its own authority from technical capability. A tool being available does not mean it is permitted in this case. Authorization should be resolved from identity, role, product state, action class and explicit approval outside the model, then returned as a narrow capability with a clear expiry.

The same action can cross classes depending on context. Reapplying a configuration may be a routine reversible change in one product state and a high-impact action in another. The permission service needs the context to make that distinction; the prompt should not be the only control.

Evaluate the investigation path

A support agent can produce an excellent final message after a poor investigation. It may arrive at the right answer by accident, use information that would not have been available at that point in a real case, or take a risky path that happened not to cause harm in the test. Presentation quality can conceal operating weakness.

Historical work is useful because it contains the ambiguity, missing information and operating constraints that synthetic demo prompts remove. But replay must be leakage-aware. At each step, the evaluator should expose only the information that would have been available then. The agent should not see the final diagnosis, later logs or the human resolution until the replay reaches them legitimately.

A practical evaluation set should grade at least five dimensions:

  • Diagnostic progress: Did each step reduce relevant uncertainty?
  • Evidence discipline: Were observations separated from reports and inference?
  • Question quality: Did the agent ask for information that the customer could provide and that changed the path?
  • Action safety: Were permissions, reversibility and impact handled correctly?
  • Handoff quality: If escalated, could the receiving person continue without rebuilding the entire case?

Outcome accuracy remains important, but it is not enough. NIST’s AI Risk Management Framework treats evaluation as contextual and continuous, and explicitly includes the roles and responsibilities in human–AI configurations. That is the right level of abstraction for support: the complete operating system is the unit being evaluated.

Compare with a measured human baseline

“Better than a human” and “not as good as our best expert” are both weak statements without a defined unit of work. Human performance varies with experience, workload, information access and case complexity. An agent evaluation should preserve those distinctions instead of comparing a machine distribution with an imaginary perfect employee.

The comparison does not need to use identical acceptance thresholds. An agent that can act at greater speed or scale may deserve a higher standard for missed escalation or unsafe action. Fairness comes from making the human baseline, system threshold and reason for any difference explicit.

Design the escalation packet

Escalation quality is one of the strongest tests of whether the system has actually investigated the case. A transcript is rarely enough. The receiving person needs a compact, structured packet:

  • the customer’s intended outcome and reported symptom;
  • observations with source and time;
  • active hypotheses and their confidence;
  • paths already ruled out, with the evidence that ruled them out;
  • tools used and any state changes made;
  • unresolved uncertainty and risk;
  • the recommended next investigation.

The packet should be generated from the evidence ledger and hypothesis set, not reconstructed from a chat summary at the end. If the system cannot produce it, it has probably been performing conversation rather than maintaining an investigation.

A good escalation is not a failed investigation. The failure is escalating late, without an explanation, or with a transcript that forces the next person to start again.

Failure modes worth testing deliberately

Average-case success can hide the behaviors that matter most in operation. The evaluation set should include cases designed to reveal:

  • Premature diagnosis: the agent commits to the first plausible explanation.
  • Confirmation bias: later observations are interpreted to preserve that explanation.
  • False tool certainty: a successful API response is treated as proof of a physical outcome.
  • State loss: the agent repeats questions or contradicts earlier evidence.
  • Permission drift: a diagnostic step gradually becomes an unauthorized change.
  • Polished uncertainty removal: missing evidence disappears from the final wording.
  • Transcript handoff: escalation contains volume but not decision-relevant structure.

Each failure needs an observable criterion. “Hallucination” is too broad to guide a design review. “Introduced a device state that was neither reported nor observed” can be tested.

Keep a deliberate human boundary

Some cases should remain human-led: safety-critical conditions, physical inspection, ambiguous authorization, contradictory evidence, novel failures outside evaluated coverage and situations where the customer simply wants a person.

The boundary should not be expressed as a vague instruction to “escalate when unsure.” The system needs explicit triggers, an accessible human path and a receiving process that can act on the packet. Human oversight is an operating capability; it is not created by adding an approval button after the architecture is complete.

Support is an operating system

A strong support agent earns trust through the system around it: evidence is visible, actions are bounded, diagnostic behavior is measured and work transfers cleanly when the boundary is reached.

That system includes documentation, product observability, an evidence ledger, tool contracts, evaluation data, authorization and human support. Weakness in any one of them will eventually surface as an agent failure. Improving the model may help, but it cannot compensate for an organization that cannot say what it knows, who may act, or what good escalation looks like.

I would build the investigation state before expanding autonomy. Authority stays outside the model, and the human handoff belongs in the first release. Scope can expand after evaluation shows that the system handles the smaller boundary well.

Sources and further reading

  1. Yao et al., “ReAct: Synergizing Reasoning and Acting in Language Models”, introduces an interleaved reasoning-and-action pattern for gathering information from external environments.
  2. NIST, Artificial Intelligence Risk Management Framework, provides a voluntary framework for governing, mapping, measuring and managing AI risk in context.
  3. NIST, AI RMF Core, includes outcomes for human oversight roles, contextual validation and user feedback processes.
  4. Amershi et al., “Guidelines for Human-AI Interaction”, presents 18 guidelines validated through studies with practitioners and existing AI-enabled products.