← Back to blogYouTube Video

Published October 2, 2026

Ai Agents Are Overrated (Here's Why)

Can't play the video or having issues? Here's the direct link.

Description

Book a call: https://calendly.com/itshassanaziz/discuss-a-project ==== ==== ==== WHAT'S THIS VIDEO ABOUT? AI agents keep creating more work than they save because they hallucinate, make wrong tool calls, need constant babysitting, and cost a lot in tokens. Chaining agents makes it worse: at 90% reliability each, a 7-agent workflow fails about half the time. I walk through how I build automations for myself and my clients instead, using an LLM once to write deterministic code (Python scripts, mostly) and calling an LLM only for the small steps that need real reasoning. Like, subscribe, & leave me a comment if you have a specific request. Thanks. ==== ==== ==== TIMESTAMPS 00:00 Why AI Agents Are Overrated 01:05 Hallucinations Break Automations 02:08 Wrong Tool Calls 02:44 Constant Babysitting 03:49 Expensive API Costs 04:51 Multi-Agent Orchestration Problem 06:06 Reliability Math: The Coin Flip 07:31 What Reliable Automations Look Like 09:19 Deterministic Code Over LLMs 10:25 My Automation Process 11:01 Benefit 1: Lower Token Spend 12:10 Benefit 2: Own Your Tooling 12:55 Benefit 3: Easy Maintenance With Tests 13:37 Where LLMs Actually Belong 14:42 From Ambiguous Steps to Deterministic Code 15:48 The 80/20 Rule of Automation 16:21 Final Takeaway WATCH THESE NEXT https://www.youtube.com/watch?v=OmV52jkTnjE https://www.youtube.com/watch?v=O8bxVFDcsIE https://www.youtube.com/watch?v=2TVhf3UvhBw ==== ==== ==== WHY LISTEN TO ME? Hey everyone, I'm Hassan. I run an AI/Automation and Software Development agency at hassandev.me. I've built custom workflows that save clients 20+ hours every week. I've helped businesses solve CRM issues that directly led to an increase of $430k CAD in quotations. I've scaled platforms to 100,000 users and beyond. I also enjoy making content and sharing what I learn with the world. I love to yap on YouTube, as you can see. I'm very active here, and on X, so if you want to reach out to me, leave a comment or DM me on X. MY LINKS Website: https://www.hassandev.me Portfolio: https://www.hassandev.me/work YouTube: https://www.youtube.com/@itshassanaziz?sub_confirmation=1 My Book: https://www.hassandev.me/designing-websites X / Twitter: https://x.com/intent/user?screen_name=nothassanaziz Instagram: https://www.instagram.com/hassansdev/ LinkedIn: https://www.linkedin.com/in/hassan-aziz-web

Transcript

Auto-generated transcript
AI agents are extremely overrated. 77% of employees are saying that using AI tools has actually increased their workload, not decreased it. And obviously, since anyone can claim things on the internet, I encourage you to go Google that stat and you're going to find a very similar number over there as well. And even just beyond all of these statistics, just look at your own workload. If you're a programmer and you're building stuff, you're probably generating more code now than you were before LLMs and before AI agents. And you're probably doing a lot more work than you were before AI agents and, you know, coding agents and everything. Why is that? Because look, the primary goal of AI was always to replace all of our jobs and, you know, put us out of our misery, right? So how does using more AI equal more work for you instead of, you know, less workloads. And the big reason is AI agents just suck. They are not reliable at all. Sure, they're powerful and they can do all sorts of cool things, but they're just not fucking reliable. And that is the thing that matters in pretty much any business out there in the world. So there's five major reasons why building an AI agent for automation usually just ends up creating more work for you than you even had to begin with. Firstly, these LLM models are going to hallucinate output. I know that this has reduced a lot over the years and I know that most you know frontier models do this less and less but make no mistake hallucinations are part of the LLM architecture. Even really strong models like Astra and Fable produce hallucinations every now and then. Hallucination just basically means you know bad invalid output that basically just breaks your entire automation because you can't reliably predict it or account for it, right? People have tried many different approaches in order to eliminate hallucinations from LLMs, but they just can't do it because it's built into the architecture. LLMs by nature are non-deterministic. If you send one single prompt to like 10 different models, even if they're the exact same model, even if they're the exact same model, they're going to produce wildly different output every single time. Now, along with that, these agents can often produce the wrong tool calls, right? So you know that an AI agent is basically an LLM that has access to a bunch of tools like, you know, reading and writing files, or I don't know, sending an email or something. Often these models are going to produce the wrong tool calls. This does get better with every single new model that these companies are releasing, but it does exist. And this issue is being caused, again because of these hallucinations because the model isn't going to reliably produce the exact same output for the exact same prompt 100% of the time and these two issues feed into the fourth issue over here which is constant babysitting because these models are non-deterministic because these agents are also non-deterministic you're stuck sitting there reviewing the LLM's output constantly and just making sure it doesn't make some dumb mistake that a human probably never would have even thought about making, right? And so these tools that you build, these agents that you set up to help automate your work just end up creating more work for you because you're stuck there constantly babysitting their output. And oftentimes, they're not going to do the entire job completely reliably. They're going to make subtle mistakes here and there, and they're going to require you to clean up their mess after they're finished. It's very difficult to get an AI agent to do a task reliably from start to completion, right? What they'll often do is they'll kind of do the 80 to 90% of the task that's really easy for you to do, and then leave the rest for you to, you know, do yourself, right? And as if that wasn't enough, as if having this tool that gets the job done some of the time wasn't a big enough mess to deal with, you also have to pay some really huge API costs because LLM tokens are not cheap. They're extremely expensive and that doesn't seem to be changing anytime soon. Compared to traditional automation that we used to have before LLMs, these tokens are extremely expensive and companies are spending millions, hundreds of millions of dollars on these things. So clearly this is not the right approach If you trying to automate most tasks right There certainly scenarios where AI agents make sense but for most business processes this is just a horrible choice to pick with. Because the big thing over here in all of these different reasons, you know, aside from the huge API costs that no company should be realistically paying, is that these agents are just not reliable. And so they're not going to reduce your workload because instead of doing the work, you're now wasting time verifying the work because you just can't trust these agents to produce reliable output, right? And so they're not going to help you at all. And this is just using one AI agent. A lot of companies or a lot of workflows are built using multi-agent orchestration layers. And what that is, it's basically using multiple different agents. I don't know, let's say you're using 10 different agents. Each agent does one particular thing and the output of one agent is fed into the other agent as the input, right? Like say you could have a workflow that is basically used to reach out to prospects or something, right? You could have one agent research the prospect. You could have the other agent draft some sort of a message or an email for the prospect. You could then feed that output into the third agent, which would then turn that draft into a proper professional email and stuff. And you could have a whole chain of AI agents running this way. The problem is, what happens if one of those agents fails, right? Because look, if the second agent fails, if let's say agent one succeeds and produces reliable output, but agent two produces bad output, that bad output gets fed into agent three, and then the outputs of agent three are going to be fed into agent four, and so on and so forth. And this is just going to cause a cascade of failures, because agent two produced bad outputs. And so just as some fun little math to do. Let's say we have a workflow that has like 10 agents, right? What's going to happen? Assuming we have 90% reliability per agent, right? This is just some really simple math. Just assume that every single agent out of these 10 is reliable 90% of the time. That means every 10 times you run it, it's going to produce a bad output one of those 10 times, right? And so if you only have one particular agent like this, it's going to produce really good output 90% of the time. If you then have two agents that work together, meaning one agent produces some output and then that gets used as the input for the second agent, the total reliability of the system just went down to 81%. If you have three agents, that goes down to 72%. You keep adding more and more agents and by the time you have like seven agents, your reliability gets reduced below 50% as in every two times you run it it's going to fail like by the time you add like six or seven agents yes I said the thing your agent workflow its reliability just got reduced to a fucking coin flip you could literally flip a coin and it would tell you whether this is going to fail or you know succeed and so as fancy as this technology is the more you go deeper and deeper into it, the less reliable it actually becomes. And the more you're doing it just as a hobby and not as some sort of like business economic value driver, right? Now, if we just discard AI agents for a second and just talk about the best automations that I've personally built, both for myself and for my clients, it's usually something like this. It's not some AI agent, but it is some AI assisted workflow, right? That is mostly deterministic, but it has an LLM step for some sort of complex reasoning that needs to occur. And it's also deterministic scripts that are just guaranteed to run the exact same way every single time you run them because they're cheap, deterministic, traditional code. And this is the kind of stuff that actually turns into a reliable automation because you can actually trust it to run the same way every single time. And it doesn't really require you to sit there verify its output. Along with that, the best automations don't require constant context switching, right? I shouldn't have to constantly monitor like 10 different automation workflows that I running I should just be able to focus on my own work And these automations should just run from the background without necessarily needing my inputs constantly because if i have to juggle like 10 different agents each producing their own output and i have to go into every single conversation thread and manage the agent and give it replies and give it answers and everything that's going to completely rot my brain and it's not going to produce any good reasonable output as you go through all three of these different automation concepts you start to think well if agents aren't going to work then the only other option we have left is deterministic code right of course that deterministic code can call llms every now and then to just execute some tiny you know step where they have to do some complex reasoning but the majority of these ideas point towards deterministic code right and that is exactly the point of this video i want to talk to you about why using traditional code oops traditional code for automation is just better than using llms and to be clear when i say that you should be using deterministic code more i don't mean handwriting the code i still want you to use llms to generate the code but it is the code that runs the automation right you'll have some sort of a python script that automates some repetitive task for you the code for that script is completely generated by an llm but it only needs to be generated once right and once it is generated you can run it as many times as you want and it's going to run the exact same way every single time that is the point over here and so that's the kind of automation that i usually build for myself and for my clients because honestly that's the point of automation right you want something that runs reliably you want something that that gets the job done without you necessarily having to think about it or you know just verify it constantly because if you're stuck doing all of that you're not really automating anything. You're just replacing one kind of work with another. You're not saving any time, any energy at all, right? And so this is the process that I follow, right? We'll use a coding agent to generate a bunch of deterministic code. That deterministic code then turns into the automations that we run. These automations run 100% reliably because these are just traditional deterministic code. It has a 0% chance of hallucinations or whatever, and we get consistent results every single time. So we don't need to sit there and verify the output of this automation because we can always trust it to perform the exact same way. And here's why this is so much better than using an AI agent that just burns tokens and isn't even reliable 100% of the time. This approach will dramatically reduce your token spend. As opposed to spending expensive LLM tokens every single time you send a message to your AI agent and ask it to do some work, you could use a coding agent to build this deterministic automation workflow for just a tiny one-time cost just for generating the code, right? And once that code is ready, you can run it like 10,000 times if you want to. It's not going to cost you any more in LLM tokens unless you have a couple steps where LLMs are actually necessary for complex reasoning. And even in those steps, since you're not using an agentic architecture, your total LLM spend is going to be vastly reduced because instead of using an agent where an LLM does everything, you're replacing like 80 to 90% of the task with deterministic code and like maybe 10% of the task where you need an LLM API call to, you know, do some reasoning, get some answers that code itself can generate. You'll probably reduce your token spend by about 90% or even all the way up down to 100%. Secondly, this approach is going to dramatically reduce your dependencies on external providers and increase your own value because you own your own tooling, right? You don't need to rely on Cloud or OpenAI or Google Gemini or whatever you're using. You don't need to rely on these external providers because you'll have your own infrastructure, your own tools for running various automations, for doing all sorts of different tasks and you don't even need you know some really high-end hosting or whatever or some like giant amount of gpus and compute to run this because this is not llms right most of this is just deterministic code that you run on your own infrastructure and you don't even need all that much like server power to run this and since you using an llm to generate all the code it becomes very very easy to manage that automation workflow because the LLM can write tests that will guarantee that the automation will work exactly the same way every single time. If you find an edge case, if you find an issue or a bug or some error in the automation, the LLM can fix that and write more tests so that the problem can never come back again. If you need to change that workflow, if you need to change how it runs, what kind of data it processes, what kind of output it produces. The LLM can do that too. All you have to do is just ask it, right? And you do this whole step one time, and then you have a finished deterministic AI automation workflow that you can run as many times as you want, and it's not going to fail on you. So I think you can kind of understand the big point that I'm making in this video, which is that the reason AI agents are failing so much, not even considering the expensive LLM tokens, because yeah, you could make the argument that, you know, as this technology matures and progresses more and more, the token expenses are going to come down and down and these tokens are going to become much cheaper. But even if you just ignore that point, these tools are just way too much unreliable, right? And so the way we introduce reliability into this whole system is by adding more deterministic code. The problem is that most of you people, or at least most people who are new to AI and new to automation, they just assume that LLMs are the end-all be-all of automation. They don't realize that we've had automation tools way before LLMs as well, and those tools have been much more reliable than LLMs. And so all we're really doing over here is we're still using LLMs for what they're best at, which is like highly complex reasoning tasks, but we're moving all the other stuff to deterministic workflows that can reliably do it for us. What ends up happening, or actually what happens at most companies when you're trying to build an automation workflow, is that the steps for that workflow are pretty ambiguous, right? Your client or your boss may tell you to, hey, go automate this workflow for me, and they're going to give you a vague bunch of steps that are just not clear yet, right? Your first instinct might be to use an LLM to automate that entire workflow, and that is great as a starter, right? You can use an LLM to automate that entire workflow because the steps currently are ambiguous. But once you run that workflow like, I don't know, 5, 10, 20 times, you're going to find that most of those steps are going to become very, very clear to you. The workflow itself is going to become much clearer to you because you're going to understand, hey, what happens when you run this workflow? What are the steps involved in it, right? Once all of those steps become clear, you're going to see that most of them can become deterministic code or deterministic scripts. right? And so you'll replace as many of those steps as you can with deterministic code so that you can boost the reliability of this automation workflow that you built by probably 90%. And in my experience, like 80% of any given automation workflow can be done using deterministic code. And around 10 or 20% requires complex LLM reasoning because you actually have to think in those scenarios as opposed to just following some, you know, naive if-else-then-if statements, conditional statements, right? But most of the workflow can easily be done using deterministic code. And that's the big point that I want you to take away from this video. Before I wrap this up, because we're done with the video over here, before I wrap this up, I just want you to know the more determinism you introduce into your AI automation workflows, the more reliable they're going to be. Using AI agents for everything is not going to be reliable because of all these different reasons that we've already discussed. Using LLMs is fine when they actually make sense for highly complex reasoning tasks where you have to think about a certain problem, very, you know, certain very complicated problem. But the more determinism you introduce into your AI workflows, the more reliable they're going to be and just overall better they're going to be. Not to mention way cheaper because you're not spending all that money on expensive LLM tokens, but just way more reliable, right? So yeah, all that to say, make your agents more reliable by introducing more Python deterministic code into it. Hope that helps.

Share this article

Hassan Aziz pointing up

Dude, you’re at the bottom of my landing page.

Just book a call already if you’re that interested.

You scrolled all the way here.