[{"data":1,"prerenderedAt":413},["ShallowReactive",2],{"post-\u002Fwhy-the-same-ai-model-produces-different-results":3},{"page":4,"translations":265,"nav":273,"related":395,"random":407},{"id":5,"title":6,"body":7,"categories":234,"category":237,"changeHistory":237,"date":238,"description":239,"disclosures":240,"draft":243,"extension":244,"firstLiveAt":237,"image":245,"imageAlt":246,"kind":247,"lang":248,"meta":249,"navigation":250,"omitGermanLocalizationDisclosure":243,"path":251,"publishedAt":237,"readingTime":252,"rights":237,"seo":253,"seoTitle":254,"slug":255,"sources":237,"stem":255,"tags":256,"translationKey":262,"type":263,"updated":237,"__hash__":264},"posts\u002Fwhy-the-same-ai-model-produces-different-results.md","Why the Same AI Model Produces Different Results Across Applications",{"type":8,"value":9,"toc":225},"minimark",[10,43,46,49,52,64,71,76,79,131,139,142,146,149,152,155,164,168,171,177,181,184,193,196,199,203],[11,12,13,21],"blockquote",{},[14,15,16,17],"p",{},"💡 ",[18,19,20],"strong",{},"TL;DR: Key Takeaways",[22,23,24,31,37],"ul",{},[25,26,27,30],"li",{},[18,28,29],{},"A model name does not describe the full working capability of an AI product."," Context, tools, permissions, and workflow also shape the result.",[25,32,33,36],{},[18,34,35],{},"The same model can do different work across applications."," It may generate text in one product and complete a task using sources and software in another.",[25,38,39,42],{},[18,40,41],{},"Test products on your own work."," Completed work, corrections, cost, failure behaviour, and authority boundaries should be evaluated together.",[14,44,45],{},"Imagine a sales team asking two AI products to prepare the same customer proposal.",[14,47,48],{},"The first turns meeting notes into polished copy. The second checks current prices and the customer's history, notices that the discount requires approval, and attaches the draft to the right CRM record. The team may assume the second product uses a more capable model. Both products could be using the same one.",[14,50,51],{},"The difference may come from the context the model can see, the tools it can use, its authority to act, and the way its work is reviewed.",[14,53,54,55,63],{},"I have previously argued that companies should ",[56,57,62],"a",{"href":58,"rel":59,"target":61},"\u002Fstart-with-the-business-problem-not-the-ai-model",[60],"noopener","_blank","start with the business problem rather than an AI model",". Once the desired result and acceptance criteria are clear, it makes sense to compare models. Even then, the model is not the only thing under evaluation.",[14,65,66],{},[67,68],"img",{"alt":69,"src":70},"The same model cartridge passes through context, tools, permissions, and workflow layers before reaching a verified record","\u002Fimages\u002Finline-model-application-harness\u002Fcapability-layers.webp",[72,73,75],"h2",{"id":74},"model-application-and-harness-are-not-the-same-thing","Model, application, and harness are not the same thing",[14,77,78],{},"An AI product can be viewed as four layers:",[80,81,82,95],"table",{},[83,84,85],"thead",{},[86,87,88,92],"tr",{},[89,90,91],"th",{},"Layer",[89,93,94],{},"Its role in the proposal example",[96,97,98,107,115,123],"tbody",{},[86,99,100,104],{},[101,102,103],"td",{},"Model",[101,105,106],{},"Interprets the notes and generates the copy",[86,108,109,112],{},[101,110,111],{},"Application",[101,113,114],{},"Lets the user start the work and review the result",[86,116,117,120],{},[101,118,119],{},"Harness",[101,121,122],{},"Manages context, tools, permissions, and stopping conditions",[86,124,125,128],{},[101,126,127],{},"Business workflow",[101,129,130],{},"Defines human responsibility and the completed outcome",[14,132,133,134,138],{},"The ",[135,136,137],"em",{},"harness"," is the operating layer around the model. It determines which information the model can see, which tools it may call, what happens after a failure, and when human approval is required.",[14,140,141],{},"These boundaries are not drawn in the same place in every product. The distinction is useful because it helps identify which decision produced the difference in results.",[72,143,145],{"id":144},"do-not-mistake-the-model-for-the-whole-product","Do not mistake the model for the whole product",[14,147,148],{},"The same model can do different work in a blank chat window, a research application, and an agent system connected to software tools. One may rely only on the text supplied by the user. Another can search current sources, preserve state across steps, create records, and request approval when necessary.",[14,150,151],{},"Product-development guides likewise treat context, tool use, approvals, failure recovery, and memory as parts of the system around the model. Raw model capability and the work a product can perform are not the same thing.",[14,153,154],{},"That means a model name is useful but incomplete. Which current and company-specific information can it access? Which actions can it perform? Are permissions enforced by software? Does it stop after a failure? Who verifies the result, and against which record?",[14,156,157,158,163],{},"I discuss the information, rules, tools, and measurement that make ",[56,159,162],{"href":160,"rel":161,"target":61},"\u002Fwhat-makes-an-ai-system-specific-to-your-business",[60],"an AI system specific to a company"," in a separate article.",[72,165,167],{"id":166},"model-names-age-evaluation-questions-last","Model names age; evaluation questions last",[14,169,170],{},"A model update can change tool use, instruction following, cost, and failure behaviour. A comparison should therefore record the model version, date, and operating configuration as well as the model name.",[14,172,173],{},[67,174],{"alt":175,"src":176},"Two offset evaluation benches compare configuration and failure records inside one shared measurement frame","\u002Fimages\u002Finline-model-application-harness\u002Fevaluation-record.webp",[72,178,180],{"id":179},"a-stronger-model-cannot-complete-a-missing-system","A stronger model cannot complete a missing system",[14,182,183],{},"A more capable model may understand complex documents better, select the right tool with less guidance, or produce the same quality at a lower cost. Model choice still matters.",[14,185,186,187,192],{},"But a stronger model cannot repair a product that lacks current data, software-enforced permissions, and a source-of-truth check for completed work. A well-designed application cannot give an unsuitable model unlimited capability either. When the steps are already known, ",[56,188,191],{"href":189,"rel":190,"target":61},"\u002Fwhen-do-you-actually-need-an-ai-agent",[60],"a fixed workflow may be more reliable and economical than an agent",".",[14,194,195],{},"When evaluating an AI product, go beyond its model name. Which context, tools, permissions, and operating structure allow it to complete your work? What results have you used to verify that capability?",[14,197,198],{},"A product's real capability appears not in the name of its model, but in the result the whole system can produce repeatedly in your work.",[72,200,202],{"id":201},"further-reading","Further Reading",[22,204,205,216],{},[25,206,207,215],{},[56,208,214],{"href":209,"rel":210,"target":61,"className":212},"https:\u002F\u002Fdevelopers.openai.com\u002Fblog\u002Fcodex-as-a-platform",[211,60],"nofollow",[213],"dofollow","Codex as a platform",": Explains how Codex treats context, tool use, approvals, and failure recovery as parts of its harness. It expands on the technical distinction between model capacity and product capability.",[25,217,218,224],{},[56,219,223],{"href":220,"rel":221,"target":61,"className":222},"https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fbuilding-effective-agents",[211,60],[213],"Building effective agents",": Discusses how agent systems extend a model with information access, tools, and memory. It reflects the provider’s own product approach and should not be read as a universal architecture rule.",{"title":226,"searchDepth":227,"depth":227,"links":228},"",2,[229,230,231,232,233],{"id":74,"depth":227,"text":75},{"id":144,"depth":227,"text":145},{"id":166,"depth":227,"text":167},{"id":179,"depth":227,"text":180},{"id":201,"depth":227,"text":202},[235,236],"ai","engineering",null,"2026-09-12","Why can the same AI model perform so differently across products? The application, context, tools, permissions, and evaluation setup all shape the result.",{"aiUse":241,"aiNote":242},"ai-assisted","AI assisted with source research, draft development, and consistency checks between the two language versions. The claims each source does and does not support were recorded separately in the research notes.",false,"md","\u002Fimages\u002Fhero\u002Fsame-model-different-applications.avif","The same blue model core can dock into a basic application or a tool-equipped work system.","Analysis","en",{},true,"\u002Fwhy-the-same-ai-model-produces-different-results",4,{"title":6,"description":239},"How Applications and Agent Harnesses Change AI Results","why-the-same-ai-model-produces-different-results",[257,258,259,260,261],"ai-models","ai-applications","agent-harness","ai-evaluation","enterprise-ai","same-ai-model-different-results","post","H8EM-82EPOyeXeKYYS-3R-meIzylqDnHr0mXktvn0Z8",{"en":266,"tr":267,"de":270},{"path":251,"title":6},{"path":268,"title":269},"\u002Ftr\u002Fayni-yapay-zeka-modeli-neden-farkli-sonuc-veriyor","Aynı Yapay Zekâ Modeli Neden Farklı Uygulamalarda Farklı Sonuç Veriyor?",{"path":271,"title":272},"\u002Fde\u002Fwarum-dasselbe-ki-modell-in-anwendungen-unterschiedliche-ergebnisse-liefert","Warum dasselbe KI-Modell in Anwendungen unterschiedliche Ergebnisse liefert",{"prev":274,"next":237,"others":277,"lucky":394,"readingTime":252},{"path":275,"title":276},"\u002Fhow-ai-changes-the-experience-gap","Can a Junior Who Uses AI Well Outperform a Senior Expert?",[278,281,284,287,290,293,296,299,302,305,308,311,313,316,319,322,325,328,331,334,337,340,343,346,349,352,355,358,361,364,367,370,373,376,379,380,383,386,389,391],{"path":279,"title":280},"\u002Ffrom-asking-questions-to-delegating-work","From Asking Questions to Delegating Work: What the Codex Study Shows",{"path":282,"title":283},"\u002Fproductlog-the-platform-i-built-for-myself-first","ProductLog: The Platform I Built for Myself First",{"path":285,"title":286},"\u002Fwhen-does-enterprise-ai-need-rag","When Does Enterprise AI Actually Need RAG?",{"path":288,"title":289},"\u002Fhow-much-does-the-eu-ai-act-affect-companies-outside-the-eu","How Much Does the EU AI Act Affect Companies Outside the EU?",{"path":291,"title":292},"\u002Fbringing-english-back-into-my-working-day","Bringing English Back Into My Working Day",{"path":294,"title":295},"\u002Fwriting-with-ai-means-thinking-with-your-archive","Writing with AI Means Thinking About Your Archive, Too",{"path":297,"title":298},"\u002Fis-your-data-ready-for-ai","Is Your Data Ready for AI? Start With the Decision It Must Support",{"path":300,"title":301},"\u002Faccessing-know-how-is-not-the-same-as-creating-it","Accessing Know-How Is Not the Same as Creating It",{"path":303,"title":304},"\u002Feu-ai-act-after-risk-classification","EU AI Act: What Should You Do Once You Know the Risk Level?",{"path":306,"title":307},"\u002Fdo-you-know-how-dependent-your-company-is-on-ai","Do You Know How Dependent Your Company Is on AI?",{"path":309,"title":310},"\u002Fwhy-sales-and-other-departments-keep-clashing","Why Sales and Other Departments Keep Clashing",{"path":160,"title":312},"What Makes an AI System Specific to Your Business?",{"path":314,"title":315},"\u002Fif-ai-handles-the-execution-who-sets-the-strategy","If AI Handles the Execution, Who Sets the Strategy?",{"path":317,"title":318},"\u002Fthe-era-of-the-previous-vibe-coder-begins","The Era of the \"Previous Vibe Coder\" Begins: The Invisibility of Clean Code and the Technical Debt Bill of AI",{"path":320,"title":321},"\u002Fhow-to-do-content-pruning-a-real-world-case-study","How to Do Content Pruning: A Real-World Case Study",{"path":323,"title":324},"\u002Fturkeys-first-real-time-mystery-shopping-reporting","Turkey's First Real-Time Mystery Shopping Reporting Platform",{"path":326,"title":327},"\u002Fgoogle-generative-ai-data-ai-citation-timing","AI Visibility Dropped Before Search: The Google Data That Changed My Theory",{"path":329,"title":330},"\u002Fone-victory-several-defeats","One Victory, Several Defeats",{"path":332,"title":333},"\u002Fllms-txt-was-never-the-point","llms.txt Was Never the Point",{"path":335,"title":336},"\u002F1m-impressions-per-month-0-revenue-a-programmatic-seo-post-mortem","1M Impressions per Month, $0 Revenue: A Programmatic SEO Post-Mortem",{"path":338,"title":339},"\u002Fai-made-code-cheap-verification-is-still-expensive","AI Made Code Cheap. Verification Is Still Expensive.",{"path":341,"title":342},"\u002Fhow-i-built-a-modern-infrastructure-using-open-source-tools-and-the-power-of-cloudflare","Why I Run camiler.org on a Single VPS",{"path":344,"title":345},"\u002Fbank-account-api-integration","Integrating One Bank Is Easy. Keeping Forty Running Is Not.",{"path":347,"title":348},"\u002Fmanaging-technology-and-transforming-the-business-are-not-the-same","Managing Technology and Transforming the Business Are Not the Same Thing",{"path":350,"title":351},"\u002Fopen-publishing-ai-monthly-seo-reporting-experiment","Open Publishing and an AI-Assisted Monthly SEO Reporting Experiment",{"path":353,"title":354},"\u002Fwhen-process-automation-actually-needs-ai","When Does Process Automation Actually Need AI?",{"path":356,"title":357},"\u002Fai-value-is-not-always-more-work","AI's Value Is Not Always More Work",{"path":359,"title":360},"\u002Fai-assisted-rest-api-development","Preserving API Quality in AI-Assisted Development",{"path":362,"title":363},"\u002Fthe-threshold-collapsed-to-zero","The Threshold Collapsed: What ProductLog Taught Me About Building in Public",{"path":365,"title":366},"\u002Fai-visibility-illusion-bing-citation-share","The AI Visibility Illusion: What Bing's Citation Share Data Actually Reveals",{"path":368,"title":369},"\u002Fprotecting-organizational-memory-during-ai-transformation","Protect Organizational Memory During an AI Transformation",{"path":371,"title":372},"\u002Fpesintaksit-cash-vs-installments-a-turkish-inflation-aware-payment-comparison-tool","Cash or Installments? – The Story Behind PeşinTaksit",{"path":374,"title":375},"\u002Fthe-job-ai-wont-take-and-the-five-it-prevents","AI Is Reducing Hiring Without Layoffs",{"path":377,"title":378},"\u002Fhow-much-authority-should-ai-have","How Much Authority Should You Give an AI System?",{"path":275,"title":276},{"path":381,"title":382},"\u002Fhow-llms-identify-experts","What Makes an LLM Recommend Someone as an Expert?",{"path":384,"title":385},"\u002Fthe-ai-productivity-baseline-is-moving-faster-than-we-remember","AI Wasn’t Always This Good. We Just Got Used to It.",{"path":387,"title":388},"\u002Fkeeping-customers-happy-isnt-enough","Keeping Customers Happy Isn’t Enough. You Have to Follow Up.",{"path":189,"title":390},"When Do You Actually Need an AI Agent?",{"path":392,"title":393},"\u002Ffrom-rules-to-decisions-the-real-time-sales-intelligence-platform-we-built-at-vanity","Medical Tourism Lead Management: How We Moved from Manual Routing to a Real-Time Sales System",{"path":323,"title":324},[396,400,402,404],{"path":397,"title":398,"date":399},"\u002Fhow-can-companies-create-value-from-ai","How Can Companies Create Value From AI?","2026-09-04",{"path":160,"title":312,"date":401},"2026-09-03",{"path":297,"title":298,"date":403},"2026-08-31",{"path":405,"title":406,"date":403},"\u002Fis-your-ai-system-delivering-business-results","Is Your AI System Actually Delivering Business Results?",[408,410,412],{"path":282,"title":283,"date":409},"2026-06-23",{"path":317,"title":318,"date":411},"2026-06-27",{"path":350,"title":351,"date":399},1789167948144]