Field guide · updated 2026-08-10 · 6 min · 1,291 words
What you should never delegate to an AI assistant
Set durable boundaries around accountability, relationships, expertise, sensitive data and irreversible action — even as model capabilities improve.

Editorial position
aiassistant.sg sells integration in this category. Product mentions carry no affiliate or vendor compensation.
Review policy
Reviewed 2026-08-10. Source links below support details that may change.
The useful answer, upfront
What to carry into the decision
- Keep final authority where accountability cannot transfer, even if AI prepares the evidence.
- Do not automate moments where a person taking responsibility is the value of the interaction.
- The less visible and recoverable a plausible error is, the stronger the independent review must be.
- Never let the model grant itself more permission or serve as the only judge of its own output.
Section 01
This is not a list of things AI cannot do
Capability lists age badly. Models improve, tools become connected and yesterday’s impossible task becomes a product feature. The useful boundaries are normative and operational: work a system should not own because accountability cannot transfer, a relationship requires a person, an error is difficult to detect or recover, or the system would be judging its own authority.
An assistant may still prepare material inside these boundaries. It can find documents, summarise options, check a form for missing fields, compare a case with an approved rubric or draft talking points. Preparation can reduce cognitive load while a named person retains the decision and the responsibility to understand the evidence.
The core question is not “can the model produce an answer?” It is “who is entitled and competent to decide, how would a plausible mistake be detected, and what would repair require?” NIST’s Generative AI Profile encourages risk management across the context and lifecycle of use, not a general verdict about the model.[3]
Working diagram
Four gates before delegation
A no at any gate means keep a meaningful human checkpoint.
01
Authority
May this decision legitimately be delegated?
02
Evidence
Can the result and its sources be independently checked?
03
Recovery
Can a plausible mistake be detected and repaired in time?
04
Relationship
Would automation remove the human responsibility people are owed?
Section 02
Where accountability cannot transfer
Keep final decisions with authorised people for contracts and significant financial commitments, regulatory attestations, hiring and dismissal, disciplinary action, medical or legal judgment and consequential decisions about a person. The model can organise evidence; it cannot assume the organisation’s legal duty or the professional’s obligation.
Singapore’s PDPC guidance for personal data in AI recommendation and decision systems is relevant where identifiable people are evaluated.[2] Accountability, meaningful information about purpose, data minimisation and appropriate human involvement should be designed into the process. A person clicking “approve” without time, context or authority is not meaningful oversight.
| Sensitive domain | AI may prepare | Person must retain |
|---|---|---|
| Hiring | Organise applications against disclosed criteria and flag missing evidence | Selection, fairness review, interview judgment and final decision |
| Financial commitment | Compare stated terms and surface anomalies | Authority to commit funds, accept risk or sign |
| Legal or regulatory | Retrieve relevant material and assemble a review pack | Interpretation, professional advice, attestation and filing approval |
| Health and safety | Collect facts and show approved procedures | Diagnosis, treatment or safety-critical judgment |
| Customer exception | Summarise history and policy options | Goodwill, liability, negotiation and accountability |
Section 03
Where accountability is part of the response
Complaint escalations, service recovery after your organisation made a mistake, difficult performance feedback and commercial negotiations are not merely drafting tasks. The other person needs someone who can understand the context, make a judgment and own what happens next.
Use an assistant backstage: build a chronology, retrieve earlier commitments, identify unresolved questions and help the accountable person prepare. The actual conversation, decision and commitment should come from the person responsible for the outcome.
Section 04
Where wrongness is invisible or expensive
Some outputs look plausible even when a specialist would recognise a fatal omission. Novel legal positions, tax edge cases, medication interactions, structural engineering and security architecture require qualified independent review. Asking the same model to critique its own answer can improve a draft, but it does not create independent assurance.
A model may also produce the right answer by an unsafe route — using the wrong record, exposing unnecessary data or succeeding only after repeated tool calls. Agent evaluation should examine outcomes, process and reliability across representative tasks rather than one polished response.[5]
- The reviewer needs the original sources and intermediate assumptions, not only the conclusion.
- The reviewer must be competent and independent enough to detect the relevant failure.
- The workflow should stop when evidence is missing rather than forcing a complete-looking output.
- The time to detection must be shorter than the time in which the error becomes irreversible.
Section 05
Actions that deserve permanent or near-permanent gates
Some tool access should remain outside a general assistant even if the surrounding workflow becomes reliable. Examples include moving substantial money, changing production access, publishing regulatory claims, deleting the only copy of data, altering identity systems, entering binding contracts and sending sensitive information to a new recipient. Expose a narrower operation with deterministic validation if automation is necessary.
OpenAI’s agent guide recommends human intervention for high-risk, irreversible or sensitive actions and when an agent exceeds failure thresholds.[4] IMDA’s agentic AI framework similarly emphasises bounded autonomy, meaningful checkpoints and human accountability.[1] These checkpoints are architectural, not signs that the system has failed to be intelligent.
| Action | Default position | Possible safer role |
|---|---|---|
| Send money or accept binding terms | Human authorisation | Prepare amount, evidence and approved destination |
| Delete or irreversibly overwrite | Human confirmation plus backup | Flag candidates and prepare a reversible archive |
| Change privileged access | Separate administrative process | Prepare a least-privilege request for an authorised admin |
| Publish sensitive or regulated content | Qualified review | Draft from approved sources with citations |
| Expand the assistant’s own scope | Independent change approval | Recommend a permission with evidence and risk statement |
Section 06
The assistant must not govern itself
Never let a model be the only evaluator of its own work, the only monitor of its failures or the authority that grants its next permission. Self-critique is useful as one signal, not an independent control. Keep policies, hard limits, approvals and system monitoring outside the model-led reasoning loop.
Likewise, do not infer safety from confidence. Confidence-like language is generated output. Use source support, deterministic validations, case tests, tool results and human judgment. If the assistant cannot show why a consequential action is permitted, it should prepare a handover rather than continue.
Section 07
A delegation ledger for every workflow
Write the boundary down so it survives staff, vendor and model changes. Every meaningful action should have an owner, autonomy level, approval trigger, evidence requirement, prohibited cases, failure route and review date. The ledger belongs with the operating workflow rather than in a presentation deck.
Step 01
Name the action
Use a concrete verb and object: send quote, update stage, cancel booking.
Step 02
Describe the consequence
Record who can be affected, reversibility and time to detection.
Step 03
Set the autonomy level
Advise, prepare, act under a rule or remain prohibited.
Step 04
Define the checkpoint
Name the authorised reviewer and the evidence they receive.
Step 05
Test the exception
Simulate missing data, conflicting policy, malicious input and tool failure.
Step 06
Review the permission
Re-authorise after material workflow, data, provider or model changes.
Section 08
The one-question test
Ask: if the assistant makes a plausible mistake here, how will we notice and repair it? If the answer is “the person affected will eventually tell us”, the workflow needs a stronger checkpoint or should not be delegated. If repair is impossible or trust itself is the asset, keep a person in the action.
Our package pages name proposed unattended actions and actions that stay human, while the final boundary is fixed in the written scope. That is a better buying artefact than a general promise that the assistant remains “human in the loop”.
↗Primary sources
Sources and verification
Citations in the article point to these first-party or authoritative references. Product details can change; the review date above is the verification date for this edition.
→Continue the field guide