Field guide · updated 2026-08-10 · 7 min · 1,351 words
How AI assistants work: models, memory, knowledge, tools and permissions
A technical but plain-language anatomy of AI assistants, including retrieval, tool calls, task state, permission gates, evaluation and failure recovery.

Editorial position
aiassistant.sg sells integration in this category. Product mentions carry no affiliate or vendor compensation.
Review policy
Reviewed 2026-08-10. Source links below support details that may change.
The useful answer, upfront
What to carry into the decision
- The model generates and decides; the surrounding system supplies current facts, persistent state, tools and constraints.
- Memory, retrieval and task state are different mechanisms. Combining them carelessly creates privacy and correctness problems.
- Every tool call is a software operation with authentication, validation, retry and logging requirements.
- Evaluation must cover the whole trajectory—inputs, sources, tool results, actions and handover—not only fluent final text.
Section 01
The model is one layer, not the product
A language model receives context and predicts a useful continuation: an answer, plan, structured value or tool request. Modern models can interpret mixed instructions and unstructured material, but they do not automatically know your current prices, remember an open task, possess permission to update a CRM or understand which mistake your business can tolerate.
An assistant wraps the model in an application. The application assembles relevant context, exposes approved tools, stores task state, validates outputs, applies permissions and records what happened. OpenAI’s agent guide reduces the core agent design to model, tools and instructions; Anthropic describes an augmented model with retrieval, tools and memory, then distinguishes fixed workflows from agents that dynamically choose their process.[1][2]
This distinction explains why model benchmarks only partially predict usefulness. A slightly stronger model connected to stale sources and broad write access can be a worse system than a modest model inside a clear workflow with good data and checks.
Working diagram
The assistant operating stack
Each layer answers a different question and should be testable on its own.
01
Experience
Where people or systems submit work and receive status or handover.
02
Workflow
Task state, routes, stopping rules, approvals and exception ownership.
03
Context
Instructions, retrieved knowledge, memory and live tool results.
04
Model
Language understanding, generation, reasoning and tool selection.
Section 02
Instructions define the job
Instructions specify the objective, relevant policy, expected output, tool rules and conditions that require handover. They can come from prompts, code, operating procedures or a combination. Good instructions convert a vague role into concrete actions: ask for an order number, retrieve the matching record, answer only from the approved policy, and route any exception to the service owner.
Instructions are not a security boundary by themselves. A model can misunderstand them, conflicting context can dilute them, and hostile content can attempt to redirect the workflow. Deterministic application code, authentication and permission checks should enforce material restrictions outside the model wherever possible.
Section 03
Memory, knowledge and task state
Memory personalises across interactions: preferences, standing instructions or prior context. Knowledge supplies authoritative material relevant to the current request. Task state records operational progress: who owns the item, what happened, what happens next and when it is due. They can all be stored, but they should not be treated as one undifferentiated conversation history.
Retrieval systems typically search or filter a source collection, select passages and place them in the model’s context for the current response. Retrieval does not guarantee truth. The collection may be stale, the wrong passage may be selected, two policies may conflict or the source may not answer the question. Useful systems preserve source identity and teach the workflow to refuse or hand over when support is insufficient.
Consumer products expose some memory controls directly. OpenAI, for example, documents saved memories and chat-history reference as controllable features.[3] A business implementation should make equivalent questions explicit: what is retained, for which purpose, who can see it, how long it remains, and how it is corrected or deleted.
| Context | Example | Owner | Failure to watch |
|---|---|---|---|
| Memory | Preferred report format or recurring constraint | The person or account owner | Stale or unexpectedly retained personal context |
| Knowledge | Current price list, policy or SOP | Named business source owner | Unsupported or superseded facts |
| Task state | Waiting on customer; next action Friday | Workflow owner | Duplicate action, lost owner or endless follow-up |
| Live tool result | Current stock, calendar availability or CRM stage | Source system | Timeout, partial response or stale cache |
Section 04
Tools let the model touch the world
A tool is a defined operation the application exposes to the model: search documents, retrieve a customer record, check a calendar, create a draft, update a field or send a message. The model selects the tool and supplies arguments; ordinary software authenticates, validates and executes it.
Tool descriptions matter because the model must distinguish similar operations and provide valid inputs. Anthropic reports that tool interface design and documentation can require as much attention as prompts.[2] Names, argument descriptions, allowed values and examples should make the safe path obvious.
Tools also fail like software. A request can time out, return an empty result, update one system but not another, or be repeated after an uncertain response. Design idempotent operations where possible, assign request identifiers, validate the returned state, cap retries and move unresolved cases to a person with the attempted actions attached.
| Failure | Unsafe reaction | Designed reaction |
|---|---|---|
| No result | Invent the missing fact | State the gap, try an approved alternative or hand over |
| Timeout | Retry indefinitely | Use a retry limit and preserve an unresolved item |
| Partial update | Assume completion | Check post-condition and reconcile affected systems |
| Permission denied | Seek broader credentials | Stop, record the boundary and notify the owner |
Section 05
Permissions determine the consequence
Permission design separates reading from writing and ordinary actions from sensitive ones. A staged assistant may begin by advising, then prepare work for approval, then gain specific reversible actions. Each expansion should name the source, action, scope, threshold, recovery path and evidence used to approve the change.
Do not give a model a powerful credential merely because the human user already has one. Create a service identity with the smallest practical access, restrict writable objects and operations, and make revocation simple. OWASP’s LLM risk guidance includes excessive agency, sensitive-information disclosure and prompt injection among the application risks that surround connected models.[4]
Singapore’s agentic AI framework recommends limiting agents’ autonomy and access to tools and data, meaningful checkpoints and whitelisted services.[6] Those controls are architecture decisions, not disclaimers.
Section 06
Orchestration and the agent loop
Some tasks follow a fixed path: classify, retrieve, draft, approve, send. Others require the model to choose steps based on intermediate results. In an agent loop, the model observes context, chooses an action, receives a tool result, updates its plan and continues until completion, a stopping condition or handover.
Dynamic planning helps with open-ended tasks, but it increases cost, latency and the chance that one error compounds through later steps. Anthropic recommends adding complexity only when simpler calls or fixed workflows do not suffice.[2] A business workflow with stable rules often benefits from deterministic orchestration around a model rather than maximum autonomy.
Working diagram
The controlled agent loop
The loop needs explicit stopping conditions and checkpoints, not an instruction to continue until satisfied.
↻
Observe
Read the task state and approved context.
↻
Choose
Select a tool or produce an answer within policy.
↻
Check
Validate the tool result, risk and completion condition.
↻
Stop or continue
Finish, request approval, hand over or take the next bounded step.
Section 07
Evaluation, monitoring and improvement
A single correct demo answer does not validate an assistant. Build a set of representative tasks with expected facts, allowed actions and handover conditions. Include incomplete requests, conflicting sources, hostile content and unavailable tools. Score groundedness, action correctness, refusal, handover completeness and final workflow state.
After launch, monitor corrections and exceptions rather than total messages alone. Group failures by source, rule, context, model, tool and permission. That diagnosis determines the fix: update a policy, improve intake, narrow a tool, add deterministic validation or change the autonomy boundary.
NIST frames generative AI risk management across governance, mapping, measurement and management.[5] In practical terms, the system needs an owner and a maintenance rhythm. The most useful assistant architecture is not the cleverest diagram; it is the one whose sources, tools, decisions and failures can be inspected and improved.
↗Primary sources
Sources and verification
Citations in the article point to these first-party or authoritative references. Product details can change; the review date above is the verification date for this edition.
- [1]OpenAI — a practical guide to building AI agents ↗
- [2]Anthropic — building effective agents ↗
- [3]OpenAI — ChatGPT memory and controls ↗
- [4]OWASP GenAI Security Project — Top 10 risks for LLM applications ↗
- [5]NIST — Generative AI Profile (AI 600-1) ↗
- [6]IMDA — Model AI Governance Framework for Agentic AI ↗
→Continue the field guide