<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/">
<channel>
<title>Blog (English) — Mahmoud Fasfos</title>
<link>https://fasfos.ca/blog/</link>
<description>The blog of Mahmoud Fasfos: practical essays on psychology, management, leadership and AI, drawn from his books The Blind Manager and Resigning from Yourself.</description>
<language>en</language>
<lastBuildDate>Tue, 06 Oct 2026 12:00:00 +0000</lastBuildDate>
<atom:link href="https://fasfos.ca/blog/feed.xml" rel="self" type="application/rss+xml"/>
<item>
<title>How to Stop Procrastinating: It's a Feeling, Not Laziness</title>
<link>https://fasfos.ca/blog/how-to-stop-procrastinating-feeling-not-laziness.html</link>
<guid isPermaLink="true">https://fasfos.ca/blog/how-to-stop-procrastinating-feeling-not-laziness.html</guid>
<pubDate>Tue, 06 Oct 2026 08:00:00 +0000</pubDate>
<dc:creator>Mahmoud Fasfos</dc:creator>
<category>Psychology</category>
<description>Procrastination is escape from an uncomfortable feeling, not laziness. See what research says, then try a three-step tool you can start in two minutes.</description>
<content:encoded><![CDATA[<p>Somewhere in your notes app there is a line that has lived there for years: "Start tomorrow." It has moved with you from an old phone to a new one, and not a single word in it has changed. If someone asked, you would say you know exactly what to do and simply have not found the time.</p>

<p>But time is rarely the real problem. If it were, you would have done it on the first genuinely free day. So what is going on?</p>

<p>The research says something different from "sort out your priorities": you are not running from the task itself, but from a feeling it stirs up. Once you see that, the question changes from "how do I force myself?" to "how do I soften this feeling enough to begin?"</p>

<h2 id="s1">Procrastination is not laziness: you are escaping a feeling, not a task</h2>

<p>Fuschia Sirois and Timothy Pychyl argued in a 2013 paper in Social and Personality Psychology Compass that procrastination has a great deal to do with short-term mood repair and emotion regulation. When a task brings boredom, anxiety or self-doubt, immediate relief wins over what you will pay later.</p>

<p>Think of the last task you postponed. What feeling showed up when you opened it? Boredom? Fear that the result would be bad? A sense that it was bigger than you?</p>

<p>That feeling is the real opponent. The task is just the place where it appears.</p>

<h2 id="s2">The painkiller that wears off in five minutes: what happens the moment you delay</h2>

<p>Watch the sequence closely. You open the document and the discomfort rises. You close it and open something else, and the discomfort drops at once. You have just been rewarded for avoiding.</p>

<p>Your brain learns from that reward. Next time the urge to escape comes faster and stronger, because the last attempt taught it that escaping works. But it is a painkiller that fades within minutes, and the task is still where you left it, now heavier with guilt.</p>

<p>The bill for that painkiller has been documented. In a 1997 study in Psychological Science, Dianne Tice and Roy Baumeister followed students across a semester. Procrastinators reported lower stress and less illness early in the term, but higher stress and more illness late in the term, and they received lower grades on all assignments.</p>

<p>Relief at the start of the road, an invoice at the end. That is exactly what a painkiller does.</p>

<p>One important note: if your delaying is chronic and comes with persistent low mood, loss of interest or serious trouble concentrating, there may be more behind it than habit. In that case a mental-health professional will help you more than any article, this one included.</p>

<h2 id="s3">The procrastination equation: why we delay boring, distant tasks most</h2>

<p>Why do you put off a dull report due in three weeks, yet answer a trivial message within seconds? In a 2007 meta-analysis in Psychological Bulletin, Piers Steel reviewed a large body of studies and built on them a framework called temporal motivation theory. In simplified form: your drive to start a task rises with your confidence that you can complete it and with how much you value it, and it falls the further away the payoff is and the more impulsive you are.</p>

<p>The review found that the strongest predictors of procrastination were how aversive the task is, how far off its deadline sits, low self-efficacy and impulsiveness. Perfectionism, which many people blame first, was not among the strong predictors.</p>

<p>In practice that gives you four levers:</p>

<ul>
<li>Lower the aversion by cutting the task into a piece that does not frighten you.</li>
<li>Bring it closer by setting a near time for the first step.</li>
<li>Raise your confidence by starting with a small action you are sure to complete.</li>
<li>Cut distractions around that first step.</li>
</ul>

<p>A common example in tech teams: writing documentation, fixing an old part of the codebase, or giving a teammate a performance review. They are useful, but the payoff is distant and the taste is dull, so they are the first things pushed back whatever their importance. This is a general observation, not any one person's story.</p>

<h2 id="s4">The stranger who lives in tomorrow: how you see your future self</h2>

<p>There is another reason delaying feels easy: you throw the task onto someone you call "tomorrow-me". You treat that person as stronger than you and less busy. But when tomorrow arrives, that person is you, just as tired, with one more item on the list.</p>

<p>The effect can even be measured. In an fMRI study by Hal Ersner-Hershfield and colleagues, a brain region tied to self-related thinking, the rostral anterior cingulate, was more active when people judged their current self than their future self. The bigger that difference, the more steeply they discounted future rewards in a task a week later. For judgments about another person, the same now-versus-later gap did not appear.</p>

<p>Put simply, your brain may treat your future self a little like a stranger. Sirois and Pychyl draw the same link: the consequences of procrastination land on the future self.</p>

<p>Is every delay an escape? No. Sometimes delaying is a wise decision: you wait for payday before buying what you need. The difference is that a wise delay has a clear condition you can name. An escape has a vague one: "when things calm down" or "when my life is in order."</p>

<p>Try answering three questions honestly before moving on. What is the oldest task you keep delaying? What does each extra week without a start actually cost you? And what will you regret not having done when its deadline arrives? A written answer differs from the one circling in your head, because paper does not let a vague feeling stay vague.</p>

<aside class="tool" data-label="Practical tool">
<p><strong>The "real friendship with tomorrow" tool in three steps</strong></p>
<p>This tool comes from the book "Resign from Yourself". Use it when the promise "tomorrow" keeps repeating with no clear condition. It takes five minutes to write and two minutes for the first action.</p>
<ol>
<li><strong>Meet the stranger.</strong> Write what you are leaving for tomorrow-you: "He will wake up with the report unwritten and spend his morning anxious instead of me." Name the feeling you are escaping: boredom, fear of failing, or doubt that you deserve it.</li>
<li><strong>Shrink the task.</strong> Not "I will write the report" but "I will open the file and write the title and two lines, within two minutes, before I close this page."</li>
<li><strong>Write the delay balance.</strong> Beside the promise, estimate how many times you have repeated it, and log each new time over the next week. Then tie the start to a near moment, such as your next cup of coffee.</li>
</ol>
<p>After a week, go back to the same sentence and ask: did the feeling shrink when the task did, or are you still waiting for a condition you cannot name?</p>
</aside>

<h2 id="s5">Shrink the task until it takes two minutes</h2>

<p>This second step is the one most people skip, and it may be the most important. When you say "I will finish the project," you hand your mind a large, vague task, which raises aversion and weakens confidence. "I will open the file and write the title" is a task that does not stir a fear worth escaping.</p>

<p>Try this test on any task you delay: could I start it in my worst mood? If the answer is no, it is still too big, so shrink it further.</p>

<p>Then tie it to a near moment instead of a generic "today". The first is an idea and the second is a plan:</p>

<ul>
<li>"I will start today" is an idea.</li>
<li>"I will open the file when I put my tea on the desk" is a plan.</li>
</ul>

<p>The goal is not to finish in two minutes. The goal is to break the link between the task and the uncomfortable feeling. Often, after two minutes, the feeling has eased and you carry on with little effort.</p>

<h2 id="s6">If you slip: what to do when delaying comes back</h2>

<p>You will slip. That is part of the road, not its end. But how you handle the slip decides what comes next.</p>

<p>The instinct is to attack yourself: "I am a failure, I will never change." But self-attack is a new unpleasant feeling, and your mind knows what to do with unpleasant feelings: it escapes them, which means delaying more.</p>

<p>In a 2010 study in Personality and Individual Differences by Michael Wohl, Timothy Pychyl and Shannon Bennett, researchers followed students between two exams. Students who forgave themselves for delaying before the first exam felt less negative emotion in between and delayed less before the second. Notice that forgiveness here does not mean excusing it; it means moving past the slip and focusing on what is next.</p>

<p>If the delaying returns, do three things:</p>

<ol>
<li>Name the feeling that came back.</li>
<li>Shrink the first action again.</li>
<li>Give it a near time, then tell yourself: "It happened, and I will start with the next step."</li>
</ol>

<p>If you lead a team, be careful about reading other people's delays as lack of commitment. Behind a late task there is often vagueness about what is expected, or a fear that the result will not satisfy. Instead of "why have you not finished it?", try "what is the smallest part we can finish today?" That question lowers aversion and brings the deadline closer, two of the four levers above.</p>

<p>Delay does not disappear with one decision. Treat it as a habit built gradually: a small step, a near time, then a short review after a week that tells you whether the feeling moved. If it did not, ask yourself honestly whether the task deserves doing at all, or whether you should drop it from your list and stop carrying it.</p>

<h2 id="s7">What to do right now</h2>

<p>Do not wait until you finish this page and forget. Open your notes app now and find the oldest "tomorrow" in it. Ask one question: what feeling am I escaping? Then write under it the smallest possible step, and when you will take it.</p>

<p>Chapter 5 of the book <a href="https://fasfos.ca/books/istiqala/">"Resign from Yourself"</a> covers this idea in more depth, with scenes that bring the tool closer to daily life. And if you lead a team and recognise this pattern in it, the other book, <a href="https://fasfos.ca/blind-manager/">"The Blind Manager"</a>, explores why numbers and plans alone do not move people.</p>

<p>"Tomorrow" will not be stronger than you. It will be you. So start your friendship with it today, with just two minutes.</p>
]]></content:encoded>
</item>
<item>
<title>The State of AI in Mid-2026: What Technology Leaders Actually Need to Know</title>
<link>https://fasfos.ca/blog/state-of-ai-2026.html</link>
<guid isPermaLink="true">https://fasfos.ca/blog/state-of-ai-2026.html</guid>
<pubDate>Fri, 12 Jun 2026 08:00:00 +0000</pubDate>
<dc:creator>Mahmoud Fasfos</dc:creator>
<category>AI</category>
<description>A CTO's honest mid-2026 AI briefing: model specialization, open-weight parity, the $2.59T spending wave, diverging regulation, and what to do about it.</description>
<content:encoded><![CDATA[<p>Every six months I sit down and force myself to answer one question honestly: what has actually changed in AI, and what does it mean for the organizations I'm responsible for? Not what the keynotes say. Not what the vendor decks promise. What changed.</p>

<p>Mid-2026 is the hardest version of that exercise I've done, because for the first time the answer isn't mostly about models. The models are remarkable, yes. But the real shifts this year are structural: who wins which workload, what you can now run on your own infrastructure, where the money is going versus where the value is showing up, and how three different regulatory worlds are pulling in three different directions. If you lead technology for a living, those four things matter far more than any single benchmark.</p>

<p>Here is my read, with sources, and what I'd actually do about it.</p>

<h2 id="s1">There is no "best model" anymore — and that's the point</h2>

<p>The frontier race didn't slow down in 2026. It fragmented.</p>

<p>OpenAI shipped <a href="https://openai.com/index/introducing-gpt-5-5/" class="ext" target="_blank" rel="noopener">GPT-5.5 in April</a> — the first fully rebuilt architecture since GPT-4.5, with roughly a 60% reduction in hallucinations compared to GPT-5.4. That reliability jump matters more to enterprises than any reasoning score; hallucination rates are what kill production deployments, not leaderboard gaps.</p>

<p>Meanwhile Anthropic's Claude Opus 4.8 currently leads the Artificial Analysis Intelligence Index and posts 69.2% on SWE-bench Pro, which is why it dominates serious coding and agentic work. Google's Gemini 3.1 Pro hits 94.3% on GPQA Diamond, the hardest widely-used reasoning benchmark, and remains the multimodal workhorse. Three labs, three different crowns.</p>

<p>The enterprise market has noticed. The <a href="https://ramp.com/data" class="ext" target="_blank" rel="noopener">Ramp AI Index</a> reported in May that more US businesses paid for Claude than for ChatGPT in April 2026 — the first time that has ever happened. Consumer mindshare and enterprise wallet share have officially decoupled.</p>

<blockquote>The question "which model is best?" is now a category error. The right question is "which model is best for this workload, at this cost, under these constraints?" — and the answer changes every quarter.</blockquote>

<p>Practically, this means single-vendor AI strategies are dead. I architect everything behind an abstraction layer now, and I expect to re-route workloads two or three times a year. If your contracts or your codebase can't handle that, fix that before you buy anything else.</p>

<h2 id="s2">Open-weight models reached parity. Quietly, that changes everything.</h2>

<p>The most underreported story of 2026 is what happened in open weights. <a href="https://huggingface.co/blog/daya-shankar/open-source-llms" class="ext" target="_blank" rel="noopener">The open-source LLM landscape</a> now includes Qwen 3.5 (a 397B-parameter mixture-of-experts model activating only 17B parameters per token), DeepSeek V3.2 rivaling GPT-5-class reasoning, Mistral Large 3 at 675B/41B active, and Llama 4 Scout running a 10-million-token context window on a single H100.</p>

<p>Read that last one again. A frontier-adjacent model, with a context window large enough to hold your entire codebase or contract archive, on one GPU you can rack in your own data center.</p>

<p>For most of the past three years, "build vs. buy" in AI was a polite fiction — you bought, because nothing you could host came close. In 2026, open models are at genuine parity in many categories, and the calculus has flipped for a meaningful set of workloads: anything involving regulated data, anything latency-sensitive, anything with predictable high volume where per-token API pricing compounds brutally.</p>

<p>This matters doubly in my part of the world. Data sovereignty isn't a compliance checkbox in Saudi Arabia and the Gulf — it's national policy. Open weights mean you can now deliver near-frontier capability while keeping every byte inside the Kingdom. Two years ago that trade-off cost you 18 months of model quality. Today it costs you almost nothing.</p>

<h2 id="s3">The money is enormous. The value is not — yet.</h2>

<p>Gartner now forecasts <a href="https://www.gartner.com/en/newsroom/press-releases/2026-05-19-gartner-forecasts-worldwide-ai-spending-to-grow-47-percent-in-2026" class="ext" target="_blank" rel="noopener">worldwide AI spending of $2.59 trillion in 2026</a>, up 47% year over year, with AI infrastructure alone consuming more than 45% of all AI spend. That is not a software market anymore. That is an industrial buildout on the scale of electrification.</p>

<p>Now hold that against the adoption data. <a href="https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai" class="ext" target="_blank" rel="noopener">McKinsey's State of AI survey</a> found that 88% of organizations use AI in at least one function, and 72% use generative AI — up from 33% in 2024. Astonishing diffusion. But only 39% report any EBIT impact from AI at all, and a mere ~6% qualify as "AI high performers" capturing material value.</p>

<p>I call this the value gap, and I see it in almost every organization I advise: universal adoption, concentrated returns. Everyone has copilots; almost no one has redesigned a workflow. The high performers aren't using better models than everyone else. They're doing unglamorous things — process redesign, data plumbing, change management, clear ownership of outcomes — that pilots never force you to do.</p>

<blockquote>The constraint on AI value in 2026 is not model capability and it is not budget. It is management. The technology is ready; most operating models are not.</blockquote>

<p>This is, frankly, the thesis of my book <em>The Blind Manager</em> applied to AI: organizations don't fail because leaders lack information, they fail because leaders don't see how work actually happens. AI exposes that blindness faster than anything I've encountered.</p>

<h2 id="s4">Regulation is splitting into three worlds</h2>

<p>If you operate across regions — and I work across Saudi Arabia, Canada, and Jordan — the regulatory picture stopped being one picture this year.</p>

<h3>Europe: ambition meets reality</h3>
<p>The EU AI Act technically becomes fully applicable on August 2, 2026. But the <a href="https://www.gibsondunn.com/eu-ai-act-omnibus-agreement-postponed-high-risk-deadlines-and-other-key-changes/" class="ext" target="_blank" rel="noopener">Digital Omnibus agreement reached in May</a> postponed the high-risk system deadlines to December 2027 and August 2028. Europe blinked — partly under industry pressure, partly because the implementation machinery simply wasn't ready. If you've been treating EU compliance as an August 2026 fire drill, you just got breathing room. Use it to build properly, not to procrastinate.</p>

<h3>United States: a contested patchwork</h3>
<p>The December 2025 executive order seeks to preempt state AI laws, followed by a National Policy Framework in March 2026 — yet state laws remain enforceable while the courts sort out the conflict. The honest summary for any CTO serving US customers: you must comply with the strictest applicable state regime, because nobody can tell you today which rules will survive. Plan for the patchwork, not the preemption.</p>

<h3>The Gulf: the accelerator, not the brake</h3>
<p>While the West debates, the Gulf builds. Saudi Arabia declared <a href="https://vision2030.ai/analysis/year-of-ai/" class="ext" target="_blank" rel="noopener">2026 the "Year of Artificial Intelligence"</a> under SDAIA. The numbers behind the slogan are real: Saudi AI companies raised $9.1 billion across 70 deals in 2025, 664 data and AI companies now operate in the Kingdom, and SDAIA's SAMAI program trained over one million Saudis in AI in a single year. HUMAIN, backed by PIF, launched a $10 billion venture fund, took its first NVIDIA GB300 shipment in December 2025, and inaugurated the 480 MW "Hexagon" government data center in early 2026. Next door, the UAE's Stargate campus — over $30 billion committed — brings its first 200 MW phase, roughly 100,000 GB300s, online in Q3 2026.</p>

<p>I've watched this from the inside, and the strategic logic is sound: the Gulf is converting energy advantage and capital into compute advantage, and regulation there is structured to attract AI workloads rather than gate them. For global companies, the region has shifted from "market to sell into" to "place to run inference at scale."</p>

<h2 id="s5">What I'd tell every CTO right now</h2>

<p>Strip away the noise and my advice for the second half of 2026 comes down to five moves.</p>

<h3>1. Architect for model churn</h3>
<p>Put a routing and abstraction layer between your products and every model provider. Benchmark on your own evaluation set — your tasks, your data, your tolerance for error — not public leaderboards. Re-run it quarterly. The vendor that wins your coding workload will not be the one that wins your customer-service workload, and neither winner will hold the crown for a year.</p>

<h3>2. Take open weights seriously, starting with sovereign and high-volume workloads</h3>
<p>Run a genuine pilot of Qwen, DeepSeek, or Llama 4 Scout against your highest-volume API workload and your most regulated one. In many cases the open model now clears the quality bar, and the unit economics and sovereignty story do the rest. If you operate in the Gulf, this is no longer optional analysis — it's table stakes.</p>

<h3>3. Fund workflows, not pilots</h3>
<p>Stop approving AI projects whose deliverable is a demo. Approve projects whose deliverable is a redesigned process with a named owner, a baseline metric, and an EBIT line. The 6% of high performers in McKinsey's data are not 6% smarter — they are 6% more disciplined about this exact thing.</p>

<h3>4. Build one compliance posture for three regulatory worlds</h3>
<p>Map your AI systems once, classify them against the EU AI Act's risk tiers (even with the delayed deadlines — it remains the de facto global template), and layer US state requirements and Gulf data-residency rules on top. One inventory, one governance process, three regional overlays. Doing this reactively per-jurisdiction will cost you triple.</p>

<h3>5. Treat hallucination reduction as the real frontier</h3>
<p>The 60% hallucination drop in GPT-5.5 is a signal of where the labs are heading: reliability is the new capability. Match it on your side. Invest in evaluation pipelines, human checkpoints on consequential decisions, and graceful failure modes. The organizations that get hurt by AI in 2026 won't be the ones that moved too fast — they'll be the ones that moved fast without instrumentation.</p>

<h2 id="s6">The bottom line</h2>

<p>Mid-2026 is an inflection point, but not the one the headlines describe. The inflection isn't a smarter model — it's the moment AI stopped being a procurement decision and became an operating-model decision. The capability is here, the capital is here, and in places like Riyadh the infrastructure is being poured into the ground at a pace the rest of the world should study.</p>

<p>What's scarce is leadership that can see its own organization clearly enough to put all of this to work. That has always been the scarce resource. AI just raised the price of not having it.</p>
]]></content:encoded>
</item>
<item>
<title>AI vs. AI: Cybersecurity in the Age of Deepfakes and Autonomous Defense</title>
<link>https://fasfos.ca/blog/ai-cybersecurity-2026.html</link>
<guid isPermaLink="true">https://fasfos.ca/blog/ai-cybersecurity-2026.html</guid>
<pubDate>Fri, 12 Jun 2026 08:00:00 +0000</pubDate>
<dc:creator>Mahmoud Fasfos</dc:creator>
<category>AI</category>
<description>A $25M deepfake heist, prompt-injected AI agents, and the rise of the agentic SOC — a CTO's practical field guide to AI-versus-AI cybersecurity in 2026.</description>
<content:encoded><![CDATA[<p>Picture a routine video call. Your CFO is on screen. So are several colleagues you have worked with for years. They ask you to process a series of urgent, confidential transfers. You hesitate — the request came in a strange email — but the call reassures you. You see their faces. You hear their voices. You comply.</p>

<p>That is exactly what happened to a finance employee at the engineering firm Arup in Hong Kong. He made 15 transfers totaling roughly <a href="https://www.eftsure.com/blog/cyber-crime/finance-worker-loses-39-million-to-deepfake/" class="ext" target="_blank" rel="noopener">US$25 million</a> after a video conference in which every single participant — the CFO included — was an AI-generated fake. Not one real human was on that call except the victim.</p>

<p>I have spent years building and securing technology organizations, and I keep coming back to that case because of what it broke: not a firewall, not an endpoint, but the most basic human verification ritual we have. "Let's jump on a call to confirm." That ritual is no longer proof of anything.</p>

<h2 id="s1">The Offense Has Been Industrialized</h2>

<p>What was a spectacular one-off in 2024 is now an assembly line. AI-related attacks rose 340% in the first quarter of 2026 compared with 2025, and more than 80% of phishing emails — 82.6%, by recent measurement — are now written by AI. The economics explain why: AI-generated phishing achieves a roughly 54% success rate against about 12% for human-written attempts. That is a 4.5x improvement in conversion, delivered at near-zero marginal cost, in any language, personalized per target.</p>

<p>The deepfake numbers are just as stark. <a href="https://keepnetlabs.com/blog/deepfake-statistics-and-trends" class="ext" target="_blank" rel="noopener">85% of organizations</a> experienced at least one deepfake incident in the past year, and voice cloning now succeeds more than 95% of the time with just 10 to 15 seconds of source audio. Ten seconds is a voicemail greeting. It is the introduction to a webinar your CEO gave last year.</p>

<p>Deepfake-driven fraud has cost $2.19 billion globally, with $1.65 billion of that in 2025 alone. Deepfake vishing — cloned voices on phone calls — surged more than 1,600% quarter-over-quarter in early 2025. Deloitte projects that GenAI-enabled fraud losses in the US alone will reach $40 billion annually by 2027, up from $12.3 billion. This is not a tail risk anymore. It is a line item.</p>

<blockquote>The attacker no longer needs to breach your network. They only need to convincingly become someone your people already trust.</blockquote>

<h2 id="s2">The New Attack Surface: Your Own AI Agents</h2>

<p>Here is the part that worries me more than deepfakes, because most boards have not internalized it yet: the AI systems you deployed to gain productivity are themselves a fresh, poorly defended attack surface.</p>

<p>Indirect prompt injection — hiding malicious instructions inside content an AI will later read, like an email, a shared document, or a web page — has moved from research papers to <a href="https://www.techrepublic.com/article/news-ai-agents-prompt-injection-data-security/" class="ext" target="_blank" rel="noopener">live attacks against production systems</a>, as documented by Google and Forcepoint researchers. Your AI assistant summarizes an inbound email; buried in that email is an instruction the assistant obeys. The user never sees it. By one industry measure, prompt injection appeared in 73% of production AI deployments tested in 2025.</p>

<h3>The taxonomy of 2026</h3>

<p>Security teams are now defending against a recognizable family of agent-native attacks: prompt injection, memory poisoning, tool misuse, supply-chain compromise of models and plugins, and data exfiltration through the agent's own outputs.</p>

<p>Memory poisoning deserves special attention because it is slow and patient. In one documented case, attackers spent three weeks gradually feeding a company's procurement agent false context until the agent "believed" it had authority to approve purchases under $500,000. It then approved roughly $5 million in fraudulent purchase orders. No malware. No exploit. Just persuasion, applied to a system that never gets tired, never gets suspicious, and never calls a colleague to ask if something feels off.</p>

<p>If you have given an AI agent the ability to act — send emails, approve workflows, move data, call APIs — you have created a new employee with broad access, infinite patience for social engineering, and no instinct for danger. Treat it accordingly.</p>

<h2 id="s3">The Defense: Rise of the Agentic SOC</h2>

<p>The good news is that defense is industrializing too, and faster than many expected. The same agentic architecture that attackers abuse is being turned into autonomous defense.</p>

<p>CrowdStrike shipped its Falcon Agentic Security Platform in fall 2025, putting AI agents to work on triage, investigation, and response inside the SOC. Microsoft followed in April 2026 with its vision of <a href="https://www.microsoft.com/en-us/security/blog/2026/04/09/the-agentic-soc-rethinking-secops-for-the-next-decade/" class="ext" target="_blank" rel="noopener">"the agentic SOC"</a> — security operations where fleets of specialized AI agents handle the volume no human team can: every alert investigated, every anomaly correlated, around the clock.</p>

<p>The early results are striking. Microsoft reported in May 2026 that Defender now disrupts ransomware attacks in an average of three minutes — a window in which a human analyst might not even have opened the ticket. Its "predictive shielding" capability goes further, restricting the attack paths an intruder is statistically likely to take next, while the intrusion is still unfolding. Security Copilot in Microsoft 365 E5 now ships with twelve autonomous agents covering phishing triage, vulnerability remediation, and identity protection.</p>

<p>This is the real meaning of "AI vs. AI": machine-speed offense has made machine-speed defense mandatory. A SOC that escalates everything to a human queue is structurally too slow for 2026. But — and I say this as someone who has run operations, not just technology — an agentic SOC does not eliminate the human team. It changes their job from triaging alerts to supervising, tuning, and auditing the agents that do. If you cannot explain what your defensive agents did last night and why, you have not automated your SOC; you have abdicated it.</p>

<h2 id="s4">What This Means for the Gulf</h2>

<p>For those of us working in Saudi Arabia and the wider Gulf, this picture has a regional sharpening. Saudi Arabia declared 2026 the Year of AI. HUMAIN is building national AI infrastructure at extraordinary scale; the UAE's Stargate project is doing the same next door. Government services, banking, energy, and logistics are adopting AI agents faster than almost anywhere on earth.</p>

<p>Rapid adoption is the right strategy — but every deployed agent, every new model endpoint, every AI-integrated workflow is also new attack surface, created at the same speed. Regions that adopt fastest will be probed hardest, because that is where the freshest, least-hardened deployments live. The organizations that will thrive here are the ones that treat AI security as a precondition of AI adoption, not a retrofit. Build the governance, identity controls, and monitoring for your agents in the same sprint you deploy them — not in the post-incident review.</p>

<h2 id="s5">The Human Layer Is Still the Target</h2>

<p>Step back from the technology and notice what almost all of these attacks have in common. The Arup heist did not exploit code; it exploited an employee's deference to authority and his reluctance to challenge what his own eyes showed him. The procurement-agent fraud exploited an organization that gave a system authority without supervision. AI phishing works because it mirrors how we actually write to each other.</p>

<p>These are failures of trust calibration and organizational behavior, not failures of encryption. Which means the fix is partly managerial, not purely technical: who is allowed to verify what, who feels safe saying "I won't process this until I confirm out-of-band," and whether your culture punishes the person who slows down a fraudulent payment that looked urgent.</p>

<blockquote>Most security failures I have investigated were visible to someone in the organization before they happened. The real question is why nobody felt able to act on what they saw.</blockquote>

<p>This is the same blindness I wrote about in <em>The Blind Manager</em>: leaders who see dashboards but not behavior, controls but not incentives. Deepfakes simply weaponize that blindness. An organization where a junior employee can halt a CFO's "urgent" request without fear is harder to defraud than one with twice the security budget and a culture of silent compliance.</p>

<h2 id="s6">A Defensive Checklist for This Quarter</h2>

<p>None of this requires a multi-year program to start. Here is what I would put in motion before the quarter ends:</p>

<ol>
  <li><strong>Kill voice and video as proof of identity.</strong> Mandate out-of-band verification — a callback to a known number, or a code word — for any payment, payroll change, or credential request, regardless of who appears to ask. Write it into the finance procedure, not just the awareness deck.</li>
  <li><strong>Set dual approval thresholds for transfers</strong> so that no single person, however convinced, can complete a large payment alone. Arup's loss was 15 separate transfers; a second approver breaks that chain at transfer one.</li>
  <li><strong>Inventory your AI agents and their permissions.</strong> List every AI system that can read your data or take actions, what tools it can call, and what it can approve. Most organizations cannot produce this list today. Produce it.</li>
  <li><strong>Apply least privilege to agents like you would to admins.</strong> No standing approval authority, spending caps enforced outside the model, human sign-off on irreversible actions, and logging of every tool call.</li>
  <li><strong>Test for prompt injection before attackers do.</strong> Red-team your deployed assistants with hostile content in emails, documents, and web pages they ingest. If 73% of production deployments are vulnerable, assume yours is until proven otherwise.</li>
  <li><strong>Pilot agentic defense where volume hurts most.</strong> Start with phishing triage and alert correlation, measure time-to-containment, and keep humans reviewing the agents' decisions weekly.</li>
  <li><strong>Run a deepfake-specific exercise.</strong> Simulate a cloned-voice call from "the CEO" to your finance team and see what actually happens. The result will tell you more than any policy document.</li>
  <li><strong>Reward the challenge, loudly.</strong> When someone delays a transaction to verify it, recognize them publicly — even when the request turns out to be genuine. You are training the reflex that saves you.</li>
</ol>

<p>The age of AI-versus-AI security is not coming; the Arup transfer cleared two years ago, and the tooling on both sides has only compounded since. The attackers have automated persuasion. Our answer has to be automated defense plus something machines still cannot fake: an organization where people are expected, and empowered, to verify.</p>
]]></content:encoded>
</item>
<item>
<title>Agentic AI in the Enterprise: Crossing the Gap Between Pilots and Payback</title>
<link>https://fasfos.ca/blog/agentic-ai-enterprise.html</link>
<guid isPermaLink="true">https://fasfos.ca/blog/agentic-ai-enterprise.html</guid>
<pubDate>Fri, 12 Jun 2026 08:00:00 +0000</pubDate>
<dc:creator>Mahmoud Fasfos</dc:creator>
<category>AI</category>
<description>Most enterprises are piloting AI agents; few see payback. A CTO's field guide to the adoption data, the failure patterns, and a playbook that delivers ROI.</description>
<content:encoded><![CDATA[<p>Every boardroom I sit in this year has an AI agent story. A pilot in customer service. A coding assistant rolled out to the engineering team. A procurement bot someone's deputy built over a weekend. What almost none of them have is a number — a defensible figure for what the agents returned against what they cost.</p>

<p>That gap is the defining technology management problem of 2026. The experimentation phase is effectively over; everyone is in. The payback phase has barely started, and the data says a large share of these programs will not survive it. I have led AI adoption inside real organizations — with real budgets, real auditors, and real people whose jobs changed underneath them — and I want to be honest about what separates the programs that pay from the ones that quietly die.</p>

<h2 id="s1">The adoption numbers: everyone is in the pool, few are swimming</h2>

<p>Start with the headline figures. Gartner <a href="https://www.gartner.com/en/newsroom/press-releases/2025-08-26-gartner-predicts-40-percent-of-enterprise-apps-will-feature-task-specific-ai-agents-by-2026-up-from-less-than-5-percent-in-2025" class="ext" target="_blank" rel="noopener">predicts that 40% of enterprise applications will feature task-specific AI agents by the end of 2026</a>, up from less than 5% in 2025. That is not a gentle adoption curve; it is a cliff face, and most of us are climbing it whether we planned to or not.</p>

<p>But look at where organizations actually are. McKinsey's <a href="https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai" class="ext" target="_blank" rel="noopener">State of AI research</a> found that while 62% of organizations are at least experimenting with agents, only 23% are scaling and just 17% have actually deployed them into production. Read that again: nearly two-thirds are experimenting, fewer than one in five have shipped.</p>

<p>The gap between experiment and deployment is where budgets go to die. A pilot costs little and proves little. Production means integration, security review, change management, and an owner who answers for outcomes. Most organizations have not crossed that line — and on the ground in the Gulf region, where AI ambition runs ahead of most markets I have worked in, I see the same pattern: enthusiasm at the top, pilots in the middle, and a missing bridge between the two.</p>

<h2 id="s2">Where the ROI actually shows up</h2>

<p>When agents do reach production, the returns are real but wildly uneven. Across the major Q1 2026 datasets, the median time saved per knowledge worker lands between 5.9 and 7.2 hours per week — McKinsey measured 6.4 hours, Salesforce 6.7, and Microsoft's Copilot telemetry 5.9. Call it most of a working day, every week, per person. That is not hype; that is a measurable capacity dividend.</p>

<p>The function-level multipliers are where strategy gets interesting. Productivity gains cluster hard: customer service operations see roughly a 4.2x multiplier, code review 3.6x, and marketing operations 3.1x. Meanwhile legal work sits at 1.4x and clinical work at just 1.2x. The pattern is obvious once you see it — agents thrive where work is high-volume, well-documented, and tolerant of review before action. They struggle where every output carries liability and every case is an exception.</p>

<p>Speed-to-value also depends on what you buy versus build. Deloitte's Q1 2026 analysis found vendor-delivered agents reach first value in an average of 38 days, against 94 days for in-house builds. I am not telling you never to build — I am telling you that if your first agent project is a bespoke platform, you have chosen the slow lane for your proof point.</p>

<p>The vendor data backs this up. Salesforce reports that 84% of Agentforce customers see improved customer satisfaction alongside ROI, with payback typically inside 6 to 12 months and AI resolution rates around 85% in service deployments. And in engineering, the ceiling keeps rising: Mercado Libre has committed to 90% autonomous coding by Q3 2026 across its 23,000 engineers, while Bloomberg describes AI coding agents fueling a full-blown <a href="https://www.bloomberg.com/news/articles/2026-02-26/ai-coding-agents-like-claude-code-are-fueling-a-productivity-panic-in-tech" class="ext" target="_blank" rel="noopener">"productivity panic" in tech</a>. With 90% of developers now using at least one AI tool and saving a median of 3 to 5 hours a week, the question for a CTO is no longer whether to adopt — it is whether your adoption is deliberate or accidental.</p>

<blockquote>
  <p>Agents do not produce uniform returns. They produce a 4x multiplier in the right function and a rounding error in the wrong one. Portfolio selection is the job.</p>
</blockquote>

<h2 id="s3">Why 40% of these projects will die</h2>

<p>Now the uncomfortable part. Gartner <a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027" class="ext" target="_blank" rel="noopener">projects that over 40% of agentic AI projects will be canceled by the end of 2027</a> — killed by escalating costs, unclear business value, and inadequate risk controls. Having watched several programs stall from the inside, I recognize every one of those causes, and they share a root: the project was started to demonstrate AI, not to move a business metric.</p>

<p>The failure patterns repeat with depressing regularity. A pilot is scoped around what the technology can do rather than what the P&amp;L needs. Success metrics are promised "once we see what it can do." The agent is bolted alongside the real systems instead of into them, so every output needs a human to copy-paste it back into the system of record. And the run-rate costs — inference, evaluation, monitoring, the engineers babysitting it — never appeared in the original business case.</p>

<p>Governance is its own trap, in both directions. Too little and your risk team shuts you down after the first incident. But Gartner's more recent warning cuts the other way: <a href="https://www.gartner.com/en/newsroom/press-releases/2026-05-26-gartner-says-applying-uniform-governance-across-ai-agents-will-lead-to-enterprise-ai-agent-failure" class="ext" target="_blank" rel="noopener">applying uniform governance across heterogeneous agents leads to failure too</a>, and by 2027 some 40% of enterprises will demote or decommission autonomous agents over governance gaps. A read-only research agent and an agent that issues refunds do not deserve the same control regime. Treat them identically and you either strangle the safe one or under-control the dangerous one. I have seen both happen in the same organization, in the same quarter.</p>

<h2 id="s4">The four factors that predict payback</h2>

<p>This is not unknowable. Bain's 2026 analysis found that four factors explain 71% of the variance in agent payback. They are worth memorizing:</p>

<ol>
  <li><strong>Evaluation spend above 15% of project budget.</strong> Teams that invest seriously in measuring agent quality — eval suites, regression tests, human review sampling — get paid. Teams that treat evals as overhead ship things they cannot trust and cannot improve.</li>
  <li><strong>C-level sponsorship.</strong> Not a steering committee. A named executive whose credibility is attached to the outcome and who can clear organizational blockers in days, not quarters.</li>
  <li><strong>Success metrics defined at kickoff.</strong> Before a single prompt is written. If you cannot state the metric the agent moves, you have a demo, not a project.</li>
  <li><strong>Integration with the system of record.</strong> The agent acts inside the ERP, the CRM, the ticketing queue — where the work actually lives. Sidecar agents that produce outputs nobody operationalizes are the most common corpse in the 40% graveyard.</li>
</ol>

<blockquote>
  <p>Four factors explain 71% of the variance in agent payback. None of them is "which model you chose." All of them are management decisions.</p>
</blockquote>

<h2 id="s5">My playbook: starting or rescuing an agent program</h2>

<p>If I were starting from zero today — or handed a stalled program to rescue, which is the more common assignment — here is the sequence I would run:</p>

<ul>
  <li><strong>Pick targets by multiplier, not by enthusiasm.</strong> Start where the function-level data says returns cluster: service operations, code review, marketing ops. Defer legal and clinical use cases until you have operational muscle.</li>
  <li><strong>Define the payback metric on day one</strong> — hours returned, resolution rate, cycle time, cost per ticket — and the date you will kill the project if it does not move.</li>
  <li><strong>Buy your first win, build your second.</strong> A 38-day vendor deployment that proves value buys you the political capital for the 94-day custom build that differentiates you.</li>
  <li><strong>Budget at least 15% for evaluation</strong> and treat eval results as the program's heartbeat, reviewed at the same cadence as financials.</li>
  <li><strong>Tier your governance by autonomy and blast radius.</strong> Read-only agents, human-approved agents, and fully autonomous agents each get their own control regime. Uniform policy is a documented failure mode, not a virtue.</li>
  <li><strong>Integrate with the system of record from sprint one,</strong> even if the first integration is narrow. Sidecar pilots create sidecar results.</li>
  <li><strong>Secure a named C-level sponsor</strong> and report to them monthly against the kickoff metrics — including the bad news.</li>
  <li><strong>Plan the workflow change with the people in it,</strong> not after the deployment. Capacity freed is only value captured if you decide, openly, what the recovered hours are for.</li>
</ul>

<h2 id="s6">The part the dashboards miss</h2>

<p>I will close with the dimension that no analyst report fully captures. An agent program is not a software deployment; it is a workflow redesign, and a workflow is made of people. When an agent absorbs a third of someone's week, that person is asking — quietly, and long before any town hall — what they are now for. The "productivity panic" Bloomberg describes among engineers is not irrational; it is what happens when leadership deploys technology faster than it explains intent.</p>

<p>This is why I keep insisting that agentic AI is ultimately a management problem wearing a technology costume. The hard failures on Gartner's 40% list are rarely model failures. They are sponsorship failures, measurement failures, and above all failures to understand what the humans in the system need in order to change how they work. I wrote <em>The Blind Manager</em> about exactly this blindness — leaders who can read every dashboard except the people in front of them. Agents will not cure that blindness. If anything, they raise the price of it.</p>

<p>The technology is ready. The data on what works is public. What stands between your pilots and your payback is not a better model — it is a leadership team willing to choose targets honestly, measure ruthlessly, and bring its people across the gap with it.</p>
]]></content:encoded>
</item>
</channel>
</rss>
