"Agentic AI" is a phrase that covers everything from a chatbot with a plugin to a system that works through a task for an hour. For engineering teams the useful question is narrower: which jobs can we hand to an agent today, and what has to be true for that to be safe? Here is how we think about it.
What an agent actually is
Strip away the marketing and an agent is a loop. A language model is given a goal and a set of tools. It decides what to do next, calls a tool, reads the result, and repeats until it is done or runs out of budget. The model supplies the judgement. Everything that makes it dependable is the engineering around it.
def run_agent(task, tools, llm, max_steps=8):
messages = [system_prompt(), user(task)]
for _ in range(max_steps):
reply = llm.complete(messages, tools=tools.schemas())
messages.append(reply)
if not reply.tool_calls:
return reply.text # finished
for call in reply.tool_calls:
result = tools.run(call) # validated, permissioned, logged
messages.append(tool_result(call.id, result))
raise StepBudgetExceeded(task) # never loop forever
Notice how much of this is ordinary software: a step budget, validated tool calls, permission checks and logging. That is where the reliability comes from.
Where agents help
Agents do best on work that is tedious, well-bounded and easy to check. In a delivery team that tends to mean:
- Test generation and repair. Drafting tests for existing code, or fixing failing ones, where the test suite itself is the judge.
- Mechanical migrations. Upgrading a framework version or replacing a deprecated API across many files.
- Triage and summarisation. Grouping incoming bug reports, or summarising a long incident thread for the next shift.
- Documentation upkeep. Keeping READMEs, runbooks and API docs in line with the code.
- First-pass code review. Catching style, obvious bugs and missing tests before a human spends time on it.
What these have in commonThere is an objective check that is not the model's own opinion: a test run, a type check, a diff a person can read in a minute.
Where they do not
- Ambiguous product decisions. An agent will produce a confident answer to a question that needed a conversation.
- Work with no way to verify the result. If you cannot tell whether it was done well, the agent cannot either.
- Irreversible actions. Deleting data, sending money or messaging customers should never happen without a person approving it.
- Architecture and trade-offs. It can suggest options. It does not carry the consequences.
Guardrails that matter
Give it the least access that works
Treat an agent like a new contractor on day one. Read-only access to the repository, a scoped token, and no production credentials. Add permissions deliberately, not by default.
Keep a human in the loop where it counts
Let the agent open a pull request, not merge one. Let it draft the incident summary, not post it. Review effort goes where the risk is.
Treat tool output as untrusted data
If an agent reads a web page, an email or a ticket, that text can contain instructions aimed at the agent. This is prompt injection, and the defence is architectural: the agent should not hold powers that an attacker could misuse through the content it reads.
Measure it with evals
You would not ship a feature without tests. Build a small set of representative tasks with known good outcomes, and run it every time you change the prompt, the model or the tools. Track success rate, cost and steps taken.
Make it observable
Log every step: what the model decided, which tool it called and what came back. When an agent does something odd, a trace is the difference between a ten-minute fix and a guess.
How we start
We suggest a small, boring first project. Pick one repetitive task your team dislikes, where success can be checked automatically. Build the narrowest agent that does it, put a person on the approval step, and measure it for a few weeks against doing the work by hand. If it earns its keep, widen the scope. If it does not, you have lost very little and learned where your real bottleneck is.
Agentic AI is a genuinely useful addition to a delivery team. It works best when it is treated as software to be engineered carefully, not as magic to be trusted.