Agentic AI: ideal employee or ticking time bomb?
What Is Agentic AI?
IBM defines it as an artificial intelligence system designed to achieve specific goals with minimal human oversight: it’s made up of AI agents, machine learning models that mimic human decision-making to solve problems in real time. In multi-agent systems, each agent handles a distinct subtask, and an orchestration layer coordinates the group toward a shared goal.

Agentic AI vs. RPA: Key Differences for Software Publishers
Traditional automation tools like RPA or Zapier follow fixed rules. Agentic AI works differently: autonomous agents replicate human decision-making instead of following a script, multi-agent orchestration splits a broader goal into subtasks handled by different agents, and contextual adaptability lets the system adjust in real time as business data changes.
LLMs at the Core of Business Processes: A Risky Revolution?
Large language models engage users so fluidly that people forget they’re talking to software, not a person. That fluency has a downside: users increasingly form something like a relationship with these agents, which opens the door to emotional harm and manipulation.
The trust problem starts with how these models actually work. They’re built on statistical analysis of enormous amounts of text, without regard for the quality or relevance of any single source, so they operate on probability rather than accuracy. In practice, that means an LLM can’t reliably tell when it doesn’t know something: pushed hard enough, it will produce a confident, “groundless” answer rather than admit the gap. And because the training data, the post-processing, the ranking algorithms, and the prompt itself are largely opaque, there’s no mathematical proof to validate what comes out the other end, which makes it hard to pin down where a given model’s limits or biases actually lie.
Solution to Our Problems, or New Problems Without Solutions?
Agentic AI’s efficiency gains are real, but so are the challenges that come with removing humans from the loop:
- Quality: with almost no human oversight, a single error or hallucination can derail an entire process, from a misassigned employee to a delayed order or a customer message that should never have gone out.
- Guarantees: classic Service Level Agreements assume deterministic systems. Once meaningful randomness enters the picture, SLAs need to be rethought from the ground up.
- Testing and validation: today’s testing frameworks weren’t built for agentic systems, and new methods are needed to define and maintain quality over time.
- Maintenance and replicability: a setup that works today isn’t guaranteed to keep working tomorrow. Prompts tuned for one model version can break on the next, forcing agents to be rewritten again and again.
- Dependence: with vendors and model versions changing fast, every update can force a partial rebuild, and that maintenance burden adds up.
- Costs and environmental impact: running agents at scale strains LLM infrastructure, and neither the computational cost nor the environmental footprint has a proven ceiling yet.
- Cyber risks: an agent that can act on your systems is also an agent that malicious actors would love to hijack, which opens new vulnerability vectors most enterprises haven’t accounted for.
Agentic AI: Avoiding the HAL 9000 Scenario
HAL 9000, from 2001: A Space Odyssey, is the classic cautionary tale: an AI built to help turns dangerous once its objectives drift from its creators’. The same risk applies, at a smaller scale, to any agent given too much autonomy and too little oversight.
That doesn’t mean giving up on agentic AI, it means treating it with more rigor. This is very concretely what we put in place at Agora on the platforms we operate:
| Strategy | Implementation Example |
|---|---|
| Continuous Control | Automated feedback loops (e.g., alerts if HR agent confidence scores drop below 90%) |
| Flexibility | Dynamic SLAs tailored to tasks (99% accuracy for orders, 90% for suggestions) |
| Resilience | Regular benchmarks and digital twins simulating worst-case scenarios |
| Sobriety | Optimize queries using lightweight models for simple tasks; measure carbon footprint per agent |

Technology alone won’t cover it, either. Cybersecurity, accountability chains, GDPR, and AI Act compliance all demand structured attention, and the human side (acceptability, workflow integration) will surface its own set of questions along the way. Most of the tools and methods needed to handle all this properly don’t exist yet, or are still in their infancy, so early deployments should expect a rough edge or two.
The Future of Agentic AI
Agentic AI still leans heavily on the pace of generative AI progress, and the limitations of today’s LLMs could well cap how far it goes in practice. Will it end up another overhyped idea left behind in the technology graveyard? Hard to say this early: ChatGPT itself is barely three years old, an eye-blink by the standards of past technology shifts, and the principles that will define mature, stable agentic systems are still being written.
Three broad scenarios seem plausible: agents become reliable enough to get woven seamlessly into everyday business functions, adoption stalls after a wave of unprofitable projects or a few high-profile incidents, or usage settles somewhere in between, concentrated in the specific domains where the technology is genuinely mature enough to trust.
None of that makes agentic AI a bad bet. It makes it an unfinished one. A single agent left unsupervised was enough to cause a 13-hour outage at a major cloud provider in 2025, as we detail in our article on MCP. The publishers who get real value out of it won’t be the ones with the flashiest demo, but the ones who pair a genuine use case with tight supervision, honest testing, and a clear view of what the model still can’t do. That discipline, more than the technology itself, is what will separate the agents still running in three years from the ones quietly switched off.
Agentic AI opens up real opportunities, but it needs a rigorous architecture to be trustworthy. The MCP protocol sits at the core of that architecture, and using it correctly makes all the difference.
Bring AI into your software with Agora Software.
Let's talk