<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community</title>
    <description>The most recent home feed on DEV Community.</description>
    <link>https://dev.to</link>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed"/>
    <language>en</language>
    <item>
      <title>EU AI Act Timeline: What AI Vendors and Developers Must Track Through 2026</title>
      <dc:creator>Ali Farhat</dc:creator>
      <pubDate>Wed, 05 Aug 2026 04:00:30 +0000</pubDate>
      <link>https://dev.to/alifar/eu-ai-act-timeline-what-ai-vendors-and-developers-must-track-through-2026-4413</link>
      <guid>https://dev.to/alifar/eu-ai-act-timeline-what-ai-vendors-and-developers-must-track-through-2026-4413</guid>
      <description>&lt;p&gt;The EU AI Act is not a single regulatory switch that turned on in 2025. Its requirements are being applied in stages, creating distinct deadlines for AI vendors, model providers, developers and organizations deploying AI systems in Europe. The first rules took effect on 2 February 2025, while a second major milestone on 2 August 2025 brought governance provisions and obligations for &lt;a href="https://scalevise.com/resources/gpai-compliance-2026-documentation-ai-office-sandbox/" rel="noopener noreferrer"&gt;providers of general-purpose AI models&lt;/a&gt; into application.&lt;/p&gt;

&lt;p&gt;That phased approach matters for companies building or using AI tooling. Consumer assistants, industrial applications and internal software may rely on general-purpose AI models, but the compliance timeline depends on the role an organization plays and the type of system involved. The next major date is 2 August 2026, when the AI Act's enforcement powers and most high-risk AI requirements are scheduled to become fully applicable.&lt;/p&gt;

&lt;h2&gt;
  
  
  The EU AI Act rollout is phased, not blanket
&lt;/h2&gt;

&lt;p&gt;The European Union has structured the AI Act around a risk-based framework and a staged implementation schedule. According to the &lt;a href="https://ai-act-service-desk.ec.europa.eu/en/ai-act/eu-ai-act-implementation-timeline" rel="noopener noreferrer"&gt;EU AI Act implementation timeline published by the AI Act Service Desk&lt;/a&gt;, 2 February 2025 marked the application of the Act's general provisions, including definitions and AI literacy, alongside prohibitions on unacceptable-risk AI uses.&lt;/p&gt;

&lt;p&gt;Six months later, on 2 August 2025, the governance framework began applying. Crucially for the current AI market, this phase also brought obligations for providers of &lt;strong&gt;general-purpose AI&lt;/strong&gt;, commonly abbreviated as GPAI, models into application. These are models with broad capabilities that can support many different downstream applications.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Implementation milestone&lt;/th&gt;
      &lt;th&gt;Date&lt;/th&gt;
      &lt;th&gt;Verified scope&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;First rules apply&lt;/td&gt;
      &lt;td&gt;2 February 2025&lt;/td&gt;
      &lt;td&gt;General provisions, including definitions and AI literacy, plus prohibitions for unacceptable-risk uses&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;GPAI and governance phase&lt;/td&gt;
      &lt;td&gt;2 August 2025&lt;/td&gt;
      &lt;td&gt;Governance framework and obligations for providers of general-purpose AI models&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Major enforcement milestone&lt;/td&gt;
      &lt;td&gt;2 August 2026&lt;/td&gt;
      &lt;td&gt;Enforcement powers and the majority of high-risk AI requirements are scheduled to become fully applicable&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Later high-risk timelines&lt;/td&gt;
      &lt;td&gt;2027 to 2028 timeframes&lt;/td&gt;
      &lt;td&gt;Further phased requirements for Annex III high-risk systems and products embedded with AI&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The distinction between these dates is essential. A company cannot reasonably reduce its AI Act planning to a single claim that the regulation either is, or is not, in force. Parts of the framework already apply, while other provisions are scheduled for later implementation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why the GPAI milestone changes the compliance conversation
&lt;/h3&gt;

&lt;p&gt;The August 2025 milestone puts particular attention on the organizations providing general-purpose AI models. This is relevant beyond the model developers themselves because many AI products and workflows are built on, or otherwise leverage, such models. For software teams, the regulatory question is therefore not limited to whether they train a foundation model. It also concerns how their product relates to the model provider, how it is deployed, and whether its intended use moves it into a more tightly regulated category.&lt;/p&gt;

&lt;p&gt;The verified timeline does not make every AI application high risk, nor does it establish identical obligations for every business that uses AI. Instead, it shows why teams need a clear view of their position in the &lt;a href="https://scalevise.com/resources/eu-ai-act-2026/" rel="noopener noreferrer"&gt;AI supply chain&lt;/a&gt; and of the systems they are placing into use.&lt;/p&gt;

&lt;p&gt;For organizations operating in Europe, the practical governance priorities supported by the rollout include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;AI literacy&lt;/strong&gt;, which entered into application with the first wave of rules in February 2025.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use-case assessment&lt;/strong&gt;, particularly to identify &lt;a href="https://scalevise.com/resources/eu-ai-act-prohibited-ai-practices/" rel="noopener noreferrer"&gt;prohibited unacceptable-risk uses&lt;/a&gt; and systems that may be subject to later high-risk requirements.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Provider and model mapping&lt;/strong&gt;, since GPAI provider obligations began applying in August 2025.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Forward planning for 2026 and beyond&lt;/strong&gt;, rather than treating the 2025 milestones as the end of implementation.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  What vendors and developers should prepare for next
&lt;/h3&gt;

&lt;p&gt;The 2 August 2026 date is the central near-term milestone for many organizations. The European Commission's phased schedule places the AI Act's enforcement powers and the majority of &lt;a href="https://scalevise.com/resources/eu-ai-act-2026-changes/" rel="noopener noreferrer"&gt;high-risk requirements&lt;/a&gt; at that point. Further timing remains relevant for Annex III high-risk systems and products that include AI, with later phases cited for 2027 and 2028.&lt;/p&gt;

&lt;p&gt;That schedule makes governance an operational issue rather than a policy exercise reserved for legal teams. Product leaders need to understand the intended use of a system. Engineering teams need clarity on which models and components are involved. Procurement and deployment decisions need to account for the provider relationships behind AI capabilities. These are practical questions that become more important as the framework expands.&lt;/p&gt;

&lt;p&gt;Companies should also avoid two unhelpful assumptions. The first is that the arrival of GPAI obligations means every tool using a general-purpose model faces the same requirements. The second is that later high-risk milestones mean current obligations can be ignored. The verified timeline supports neither conclusion. The Act is already applying in defined areas, while additional requirements are still approaching.&lt;/p&gt;

&lt;p&gt;Organizations assessing AI systems, model dependencies and internal governance can work with Scalevise on AI architecture, &lt;a href="https://scalevise.com/resources/ai-governance/" rel="noopener noreferrer"&gt;workflow automation&lt;/a&gt; and implementation planning that connects technical delivery with operational controls.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;When did the first EU AI Act rules start applying?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The first wave started on 2 February 2025. It applied general provisions, including definitions and AI literacy, as well as prohibitions for unacceptable-risk AI uses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When did obligations for general-purpose AI model providers begin?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Obligations for providers of general-purpose AI models began applying on 2 August 2025, alongside the AI Act's governance framework.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does the EU AI Act fully apply to every AI system already?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No. The Act is rolling out in stages. Enforcement powers and most high-risk AI requirements are scheduled to become fully applicable on 2 August 2026, with further high-risk timelines extending into 2027 and 2028.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why should companies using third-party AI models track the GPAI rules?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Many AI tools rely on or leverage general-purpose AI models. Companies need to understand their role in the AI supply chain, their system's intended use and the implementation dates that may apply to them.&lt;/p&gt;




&lt;h3&gt;
  
  
  Conclusion
&lt;/h3&gt;

&lt;p&gt;The EU AI Act's 2025 milestones established that AI governance in Europe is already a live operational concern, especially for unacceptable-risk uses, AI literacy and general-purpose AI model providers. For vendors and developers, the key task is to treat the regulation as a staged program: address the rules already in application, map AI dependencies and prepare for the major high-risk and enforcement milestone scheduled for August 2026.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>automation</category>
      <category>tools</category>
      <category>governance</category>
    </item>
    <item>
      <title>Finding Exposed Services (and Vulnerabilities) on Your Network: A Practical ScanSearch Guide</title>
      <dc:creator>Billy</dc:creator>
      <pubDate>Wed, 05 Aug 2026 04:00:13 +0000</pubDate>
      <link>https://dev.to/devyjones/finding-exposed-services-and-vulnerabilities-on-your-network-a-practical-scansearch-guide-20e3</link>
      <guid>https://dev.to/devyjones/finding-exposed-services-and-vulnerabilities-on-your-network-a-practical-scansearch-guide-20e3</guid>
      <description>&lt;p&gt;Ever wondered what services on your network are truly exposed to the internet? Or perhaps you're trying to track down a specific type of device across your public-facing infrastructure? Manually scanning port by port across a large IP range can be tedious, slow, and often misses the bigger picture.&lt;/p&gt;

&lt;p&gt;This article isn't about setting up a traditional port scanner. Instead, we're going to explore how to leverage an internet-wide search engine for network devices and services, &lt;a href="https://scansearch.net" rel="noopener noreferrer"&gt;ScanSearch&lt;/a&gt;, to quickly identify and understand the public-facing footprint of your network. Think of it like Google, but for servers, routers, cameras, and everything else connected to the internet.&lt;/p&gt;

&lt;p&gt;We'll cover practical use cases, including:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Identifying unintentionally exposed services.&lt;/li&gt;
&lt;li&gt;  Searching for specific device types or software versions.&lt;/li&gt;
&lt;li&gt;  Spotting known vulnerabilities associated with your infrastructure.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Let's dive in.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem: Blind Spots in Your Public-Facing Infrastructure
&lt;/h2&gt;

&lt;p&gt;It's common for development or operations teams to configure services, sometimes forgetting that a default configuration or an overlooked firewall rule might leave something accessible that shouldn't be. This isn't just about malicious actors; it's also about maintaining good security hygiene and understanding your own attack surface.&lt;/p&gt;

&lt;p&gt;Traditional internal network scans are crucial, but they don't always give you the external perspective. That's where a tool like ScanSearch comes in – it's constantly indexing the internet's public-facing devices.&lt;/p&gt;

&lt;h2&gt;
  
  
  Getting Started with ScanSearch
&lt;/h2&gt;

&lt;p&gt;ScanSearch offers a powerful search syntax. The most straightforward way to begin is by searching for your organization's public IP ranges or domain names. For this tutorial, I'll use &lt;code&gt;example.com&lt;/code&gt; and a hypothetical IP range, but you should substitute these with your own public IPs or domains.&lt;/p&gt;

&lt;h3&gt;
  
  
  Use Case 1: Discovering All Services on Your IP Range
&lt;/h3&gt;

&lt;p&gt;Let's say your organization uses the IP range &lt;code&gt;203.0.113.0/24&lt;/code&gt;. To see everything ScanSearch has indexed for this range, you'd simply enter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ip:203.0.113.0/24
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This query will return a list of all devices, services, and associated information (like open ports, banners, and even HTTP response headers) found within that IP range. You might be surprised at what pops up – maybe an old development server you thought was offline, or a service running on a non-standard port.&lt;/p&gt;

&lt;h3&gt;
  
  
  Use Case 2: Finding Specific Device Types or Software Versions
&lt;/h3&gt;

&lt;p&gt;Perhaps you're concerned about older versions of Apache or Nginx that might still be running. You can combine the IP range with a keyword search for specific banners or technologies. For example, to find all Apache servers on our hypothetical range:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ip:203.0.113.0/24 product:apache
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or, to get more specific and look for an older, potentially vulnerable version (e.g., Apache 2.2):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ip:203.0.113.0/24 product:apache/2.2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This quickly highlights potential upgrade candidates or misconfigurations.&lt;/p&gt;

&lt;h3&gt;
  
  
  Use Case 3: Identifying Known Vulnerabilities on Your Infrastructure
&lt;/h3&gt;

&lt;p&gt;ScanSearch also indexes publicly known vulnerabilities (CVEs) associated with identified services. This is incredibly powerful for proactively assessing risk.&lt;/p&gt;

&lt;p&gt;Let's say you want to see if any of your devices on &lt;code&gt;example.com&lt;/code&gt; are running services with known vulnerabilities:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;host:example.com has_vulnerability:true
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This query would return any hosts under &lt;code&gt;example.com&lt;/code&gt; where ScanSearch has identified a service with an associated, known vulnerability. You can then drill down into the results to see the specific CVEs and affected services. This is a critical step in prioritizing patching and mitigation efforts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pro-tip:&lt;/strong&gt; You can also search directly for a specific CVE. If you're tracking a new, critical vulnerability, you could search &lt;code&gt;cve:CVE-2023-XXXX&lt;/code&gt; to see if it appears anywhere on your network.&lt;/p&gt;

&lt;h3&gt;
  
  
  Use Case 4: Searching for Default Credentials or Common Misconfigurations
&lt;/h3&gt;

&lt;p&gt;While ScanSearch doesn't actively exploit systems, its indexing of service banners and HTTP responses can sometimes reveal clues about weak configurations. For instance, you might look for common default administrative interfaces or specific keywords that indicate a lack of proper setup.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ip:203.0.113.0/24 title:"admin panel" 
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or, looking for specific HTTP response bodies:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ip:203.0.113.0/24 http.body:"Powered by WordPress"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;While not a direct vulnerability, finding an exposed WordPress admin panel on a public IP range you own is certainly something you'd want to investigate further.&lt;/p&gt;

&lt;h2&gt;
  
  
  Beyond Simple Queries: Combining Filters
&lt;/h2&gt;

&lt;p&gt;The real power of ScanSearch comes from combining these filters. You can search for specific ports, protocols, HTTP headers, and much more. For example, to find all devices on your range running an SSH server (port 22) that also have a known vulnerability:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ip:203.0.113.0/24 port:22 has_vulnerability:true
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This quickly narrows down your focus to the most critical issues.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Understanding your external attack surface is a fundamental part of maintaining secure infrastructure. Tools like &lt;a href="https://scansearch.net" rel="noopener noreferrer"&gt;ScanSearch&lt;/a&gt; provide an invaluable perspective by indexing the entire internet, allowing you to quickly query and identify exposed services, potential misconfigurations, and known vulnerabilities associated with your own public-facing assets.&lt;/p&gt;

&lt;p&gt;Instead of endless manual scans, integrate ScanSearch into your regular security audits. It's a powerful way to catch those overlooked services and proactively address potential risks before they become a problem. Give it a try with your own public IP ranges and domains – you might be surprised at what you find.&lt;/p&gt;

</description>
      <category>security</category>
      <category>networking</category>
      <category>devops</category>
      <category>tooling</category>
    </item>
    <item>
      <title>Designing a Reliable PDF Translation Job Pipeline in TypeScript</title>
      <dc:creator>Noah Chen</dc:creator>
      <pubDate>Wed, 05 Aug 2026 03:57:55 +0000</pubDate>
      <link>https://dev.to/noahchenbuilds/designing-a-reliable-pdf-translation-job-pipeline-in-typescript-1i04</link>
      <guid>https://dev.to/noahchenbuilds/designing-a-reliable-pdf-translation-job-pipeline-in-typescript-1i04</guid>
      <description>&lt;p&gt;Uploading a PDF and calling a translation model looks like a two-step feature. In production, it is a job pipeline with untrusted input, two different extraction paths, several expensive stages, and an output that can be fluent while still being wrong.&lt;/p&gt;

&lt;p&gt;That distinction matters for a small SaaS team. The translation request may come from support, sales, or an internal operations task. Nobody wants to operate a document platform, but the workflow still needs to answer basic questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Was the upload actually a PDF?&lt;/li&gt;
&lt;li&gt;Does the file contain selectable text or scanned page images?&lt;/li&gt;
&lt;li&gt;Can a retry create a second charge or a conflicting result?&lt;/li&gt;
&lt;li&gt;What happens when page 37 fails after the first 36 pages succeed?&lt;/li&gt;
&lt;li&gt;How do we know the translated PDF is not blank or visually broken?&lt;/li&gt;
&lt;li&gt;When are the source and result deleted?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The translation model is one component. Reliability comes from the system around it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Define the Job Contract First
&lt;/h2&gt;

&lt;p&gt;I would not let a file reach an extractor until the API has established a narrow contract.&lt;/p&gt;

&lt;p&gt;For example, a translation request might include:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;TranslationStyle&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;general&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;technical&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;academic&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;CreateTranslationJob&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;uploadId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;sourceLanguage&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;auto&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;targetLanguage&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;style&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;TranslationStyle&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;idempotencyKey&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;containsRestrictedData&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;boolean&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The request should be rejected when the source and target languages are identical, the upload is missing, the target language is unsupported, or policy says the document cannot leave an approved environment.&lt;/p&gt;

&lt;p&gt;File validation should also be explicit. Do not trust the filename or browser-supplied MIME type. Check at least:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;the actual byte size;&lt;/li&gt;
&lt;li&gt;the file signature;&lt;/li&gt;
&lt;li&gt;whether the parser can open the document;&lt;/li&gt;
&lt;li&gt;whether the PDF is encrypted;&lt;/li&gt;
&lt;li&gt;the page count;&lt;/li&gt;
&lt;li&gt;whether the job fits the account or product limit.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A 20 MB limit is simple to explain in a user interface, but size alone is not a good predictor of work. A compressed 200-page text PDF can be smaller than a six-page scan. Page count, image area, and extracted character count are better inputs for estimating processing time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat Preflight as Its Own Stage
&lt;/h2&gt;

&lt;p&gt;The first useful result is not a translation. It is a document profile.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;DocumentProfile&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;pageCount&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;encrypted&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;boolean&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;extractedCharacters&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;pagesWithText&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;pagesWithLargeImages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;likelyScanned&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;boolean&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;estimatedOcrPages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This profile controls routing. A PDF where nearly every page has substantial selectable text can go directly to extraction. A document made of page-sized images needs OCR. Mixed documents need a page-level decision rather than a single flag for the entire file.&lt;/p&gt;

&lt;p&gt;A crude scan heuristic could look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;needsOcr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;profile&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;DocumentProfile&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nx"&gt;boolean&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;profile&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;pageCount&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;textCoverage&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;profile&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;pagesWithText&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nx"&gt;profile&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;pageCount&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;imageCoverage&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;profile&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;pagesWithLargeImages&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nx"&gt;profile&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;pageCount&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;textCoverage&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mf"&gt;0.25&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;imageCoverage&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mf"&gt;0.6&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The thresholds are product decisions, not universal constants. A form can contain a small amount of real text over a scanned background. A research paper may include image-heavy appendix pages. Store the measurements that led to the route so a failed job can be explained later.&lt;/p&gt;

&lt;h2&gt;
  
  
  Model the Pipeline as States, Not Progress Percentages
&lt;/h2&gt;

&lt;p&gt;A single &lt;code&gt;processing: true&lt;/code&gt; field hides too much.&lt;/p&gt;

&lt;p&gt;I would rather expose a finite set of states:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;JobState&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;queued&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;validating&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;extracting&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ocr&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;translating&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;rendering&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;verifying&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ready&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;failed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;expired&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each state should have a clear owner and output:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;State&lt;/th&gt;
&lt;th&gt;Responsibility&lt;/th&gt;
&lt;th&gt;Durable output&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;validating&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;inspect the upload and policy&lt;/td&gt;
&lt;td&gt;document profile&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;extracting&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;recover text and geometry&lt;/td&gt;
&lt;td&gt;page blocks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ocr&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;recognize image-only pages&lt;/td&gt;
&lt;td&gt;OCR blocks and confidence&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;translating&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;translate normalized blocks&lt;/td&gt;
&lt;td&gt;translated segments&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;rendering&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;rebuild the destination PDF&lt;/td&gt;
&lt;td&gt;candidate output&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;verifying&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;run structural checks&lt;/td&gt;
&lt;td&gt;verification report&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ready&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;expose a time-limited download&lt;/td&gt;
&lt;td&gt;signed result reference&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The state machine makes several bugs harder to create. A rendering worker cannot start before translation output exists. An expired job cannot silently return to &lt;code&gt;ready&lt;/code&gt;. A retry can resume from the last durable stage instead of repeating the entire workflow.&lt;/p&gt;

&lt;p&gt;Transitions should be enforced rather than implied:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;allowed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;Record&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;JobState&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;JobState&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;queued&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;validating&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;failed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;validating&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;extracting&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ocr&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;failed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;extracting&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ocr&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;translating&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;failed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;ocr&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;translating&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;failed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;translating&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;rendering&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;failed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;rendering&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;verifying&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;failed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;verifying&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ready&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;failed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;ready&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;expired&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;failed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[],&lt;/span&gt;
  &lt;span class="na"&gt;expired&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[],&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;canTransition&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JobState&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;to&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JobState&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nx"&gt;boolean&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;allowed&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;to&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In a real system I would also record &lt;code&gt;attempt&lt;/code&gt;, &lt;code&gt;updatedAt&lt;/code&gt;, &lt;code&gt;failureCode&lt;/code&gt;, and the worker version that produced each stage. That is enough to distinguish a bad document from a deployment regression.&lt;/p&gt;

&lt;h2&gt;
  
  
  Put Retry Boundaries Around Durable Work
&lt;/h2&gt;

&lt;p&gt;Retries are necessary, but “retry the job” is too broad.&lt;/p&gt;

&lt;p&gt;Uploading, OCR, translation, and rendering have different failure modes. A network timeout while writing the final PDF should not trigger OCR again. A transient translation-provider error should not require another upload. A user pressing the submit button twice should not create two independent jobs.&lt;/p&gt;

&lt;p&gt;The idempotency key belongs at job creation. Stage-specific keys can be derived from the job and input version:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;stageKey&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;jobId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;stage&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JobState&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;inputVersion&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;jobId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;stage&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;inputVersion&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Workers should write their result before advancing the state. If the process stops between those operations, the next worker can see the existing stage output and continue safely.&lt;/p&gt;

&lt;p&gt;Retries also need limits. OCR on a malformed image is unlikely to improve on attempt 12. Use stable failure codes such as &lt;code&gt;PDF_ENCRYPTED&lt;/code&gt;, &lt;code&gt;OCR_LOW_CONFIDENCE&lt;/code&gt;, &lt;code&gt;TRANSLATION_TIMEOUT&lt;/code&gt;, and &lt;code&gt;RENDER_OVERFLOW&lt;/code&gt;. They are more useful to support than a stack trace alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Translation Units Need Stable Identity
&lt;/h2&gt;

&lt;p&gt;Sending one entire document as a string loses layout relationships. Sending every line independently loses context.&lt;/p&gt;

&lt;p&gt;A practical middle ground is a collection of blocks with stable IDs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;TextBlock&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;page&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;heading&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;paragraph&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;caption&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;table-cell&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;footer&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;sourceText&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;translatedText&lt;/span&gt;&lt;span class="p"&gt;?:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;bounds&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;x&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;y&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;width&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;height&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="nl"&gt;ocrConfidence&lt;/span&gt;&lt;span class="p"&gt;?:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Stable IDs make it possible to retry one segment, preserve repeated headers, compare source and translation, and point a human reviewer to a specific page and region.&lt;/p&gt;

&lt;p&gt;Context can be supplied by grouping adjacent blocks or attaching a document glossary. The glossary should protect product names, UI labels, units, and phrases that must remain unchanged. It is also the right place to keep reviewer-approved terminology between versions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verification Is More Than “The File Opens”
&lt;/h2&gt;

&lt;p&gt;A translated PDF can be syntactically valid and unusable.&lt;/p&gt;

&lt;p&gt;Automated verification should look for structural signals:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;output page count differs unexpectedly from the source;&lt;/li&gt;
&lt;li&gt;pages contain no text or images;&lt;/li&gt;
&lt;li&gt;translated blocks are missing;&lt;/li&gt;
&lt;li&gt;text extends outside its assigned bounds;&lt;/li&gt;
&lt;li&gt;a table has a different number of rows or cells;&lt;/li&gt;
&lt;li&gt;a required font lacks glyphs for the target language;&lt;/li&gt;
&lt;li&gt;hyperlinks disappeared;&lt;/li&gt;
&lt;li&gt;OCR confidence is below a review threshold;&lt;/li&gt;
&lt;li&gt;numbers present in the source are absent from the translated block.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These checks do not prove semantic accuracy. They decide whether the result is safe to show as an ordinary completion or should be flagged for review.&lt;/p&gt;

&lt;p&gt;For customer-facing, legal, medical, financial, or safety-related documents, a human review is not an optional fallback. The workflow should make that requirement visible instead of letting a fluent result imply certification.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retention Is Part of the Data Model
&lt;/h2&gt;

&lt;p&gt;Document pipelines tend to become accidental archives.&lt;/p&gt;

&lt;p&gt;The source upload, extracted text, OCR images, model inputs, intermediate render, and final file may all exist in different systems. Deleting only the public download does not remove the job data.&lt;/p&gt;

&lt;p&gt;Every artifact should carry a retention class and expiry:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;StoredArtifact&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;key&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;source&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;extracted&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ocr&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;translated&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;result&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;deleteAfter&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;encrypted&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;boolean&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use short-lived signed URLs for downloads. Run deletion as an observable job. Record completion without keeping the deleted content. If business users need a permanent copy, make them move it into the approved system of record rather than silently extending temporary storage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make Operations Visible
&lt;/h2&gt;

&lt;p&gt;A progress bar is useful to the person waiting, but it is not enough for the team operating the workflow.&lt;/p&gt;

&lt;p&gt;I would track metrics by stage and route:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;validation failures by reason;&lt;/li&gt;
&lt;li&gt;percentage of pages sent to OCR;&lt;/li&gt;
&lt;li&gt;median and high-percentile duration for extraction, OCR, translation, and rendering;&lt;/li&gt;
&lt;li&gt;retries and terminal failures by worker version;&lt;/li&gt;
&lt;li&gt;output files flagged for human review;&lt;/li&gt;
&lt;li&gt;deletion jobs completed after their deadline;&lt;/li&gt;
&lt;li&gt;jobs abandoned before download.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These measurements answer different questions. A sudden increase in OCR pages may indicate a change in customer inputs. Longer rendering time after a deployment points toward layout code, not the translation provider. A high completion rate with a low download rate may mean results expire too quickly or the notification step is unreliable.&lt;/p&gt;

&lt;p&gt;Logs should carry the job ID, stage, attempt, and block or page range, but not extracted document text by default. Document contents make debugging tempting and data handling much harder. Store diagnostic metadata first; collect a redacted sample only through an explicit support path.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build, Run Locally, or Use a Hosted Workflow?
&lt;/h2&gt;

&lt;p&gt;Not every team should build this pipeline.&lt;/p&gt;

&lt;p&gt;Building makes sense when translation is part of the product, jobs repeat at scale, restricted documents are involved, or the team needs precise control over models, audit logs, and retention. Local tools are attractive when a technical operator can process sensitive files without uploading them.&lt;/p&gt;

&lt;p&gt;For an occasional public or synthetic document, a hosted interface can remove a lot of setup. &lt;a href="https://pdftranslator.org/" rel="noopener noreferrer"&gt;PDFTranslator&lt;/a&gt; is one example: it supports OCR, more than 100 languages, and translation styles, with a 20 MB upload limit and a stated free allowance of 1,000 pages per calendar month. It is still a cloud workflow. Its homepage says task files are deleted within 24 hours, which is useful operational information but not a reason to upload restricted material without an approved policy.&lt;/p&gt;

&lt;p&gt;That tradeoff is the important part. Convenience changes who operates the pipeline; it does not remove the need to classify the document or review the result.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Production-Readiness Checklist
&lt;/h2&gt;

&lt;p&gt;Before calling a PDF translation workflow reliable, I would want clear answers to these questions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Do we validate file contents rather than extensions?&lt;/li&gt;
&lt;li&gt;Can we route text, scanned, and mixed PDFs differently?&lt;/li&gt;
&lt;li&gt;Are job states and legal transitions explicit?&lt;/li&gt;
&lt;li&gt;Are retries idempotent and limited to the failed stage?&lt;/li&gt;
&lt;li&gt;Can we trace translated blocks back to source regions?&lt;/li&gt;
&lt;li&gt;Do we verify page structure, overflow, fonts, tables, links, and numbers?&lt;/li&gt;
&lt;li&gt;Can the system require human review instead of merely suggesting it?&lt;/li&gt;
&lt;li&gt;Does every stored artifact have a deletion rule?&lt;/li&gt;
&lt;li&gt;Can support explain a failure without reading worker logs?&lt;/li&gt;
&lt;li&gt;Have we decided which documents are allowed to leave the device?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The model call may be the most visible part of translation, but it is not the system. The system is the set of boundaries that keeps an untrusted document, a long-running process, and a persuasive-looking output from creating a quiet operational failure.&lt;/p&gt;

</description>
      <category>typescript</category>
      <category>architecture</category>
      <category>webdev</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>Building an Editable 3D Indoor Map in the Browser</title>
      <dc:creator>zyl19950114</dc:creator>
      <pubDate>Wed, 05 Aug 2026 03:56:10 +0000</pubDate>
      <link>https://dev.to/zyl19950114/building-an-editable-3d-indoor-map-in-the-browser-1md0</link>
      <guid>https://dev.to/zyl19950114/building-an-editable-3d-indoor-map-in-the-browser-1md0</guid>
      <description>&lt;p&gt;Indoor maps are often treated as a rendering problem: take a floor plan, extrude a few walls, and display the result. That is useful for a viewer, but it breaks down when a team needs to edit a real space, place assets, or hand the result to another application.&lt;/p&gt;

&lt;p&gt;We are building &lt;strong&gt;KiMap&lt;/strong&gt; around a different boundary: turn a floor plan into an editable indoor scene in the browser, then keep the resulting structure useful for an SDK consumer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a floor plan is not enough
&lt;/h2&gt;

&lt;p&gt;A production indoor workflow needs more than a textured image on a plane. At minimum, the editor has to preserve the relationships between walls, floors, rooms, openings, and the objects placed in the space. Those relationships determine whether the result can later support navigation, facility workflows, a digital twin, or a custom web experience.&lt;/p&gt;

&lt;p&gt;That is why the current KiMap workflow starts with structure. You can define the indoor geometry, inspect it in 2D and 3D, and keep editing instead of committing to a static export too early.&lt;/p&gt;

&lt;h2&gt;
  
  
  The browser editor boundary
&lt;/h2&gt;

&lt;p&gt;The editor is built with React and Three.js. The goal is not to replace every DCC tool. It is to make the early spatial workflow accessible to teams that need to test an indoor experience before investing in a full custom pipeline.&lt;/p&gt;

&lt;p&gt;The parts we are concentrating on are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;editable floor-plan structure and bounded spaces&lt;/li&gt;
&lt;li&gt;2D and 3D scene inspection in the same workflow&lt;/li&gt;
&lt;li&gt;reusable 3D furniture and local asset handling&lt;/li&gt;
&lt;li&gt;saving an indoor project without dropping the referenced model data&lt;/li&gt;
&lt;li&gt;a path toward SDK-oriented rendering and integration&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The last point matters. A scene that looks correct in an editor is not automatically useful to an application. We want the data boundary to be explicit enough that an SDK consumer can load the geometry and assets without rebuilding the scene from scratch.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we are testing next
&lt;/h2&gt;

&lt;p&gt;KiMap is in free early access. The most useful feedback is not generic interest; it is a concrete blocker from someone building an indoor-navigation, GIS, WebGL, facility, or digital-twin workflow.&lt;/p&gt;

&lt;p&gt;Try the live example without creating an account:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.kimap.cc/editor/example?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=early_access&amp;amp;utm_content=first_article" rel="noopener noreferrer"&gt;Open the KiMap example&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;After trying it, I would especially value answers to three questions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Which indoor data model or export format would your stack require?&lt;/li&gt;
&lt;li&gt;Where would this editing workflow conflict with your real building or facility data?&lt;/li&gt;
&lt;li&gt;What API or model-loading behavior would you need before integrating it?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Early access is free and no credit card is required. There is a feedback action inside the example, or you can reach the team at &lt;a href="mailto:kimap.founder@gmail.com"&gt;kimap.founder@gmail.com&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>threejs</category>
      <category>webdev</category>
      <category>javascriptlibraries</category>
    </item>
    <item>
      <title>Seedance 2.5 is priced 53% above 2.0 per token, and its 480p frame shrank</title>
      <dc:creator>Lee</dc:creator>
      <pubDate>Wed, 05 Aug 2026 03:55:51 +0000</pubDate>
      <link>https://dev.to/lee_315dd1e13420e63e2b813/seedance-25-is-priced-53-above-20-per-token-and-its-480p-frame-shrank-5fki</link>
      <guid>https://dev.to/lee_315dd1e13420e63e2b813/seedance-25-is-priced-53-above-20-per-token-and-its-480p-frame-shrank-5fki</guid>
      <description>&lt;p&gt;Seedance 2.5's API opens on August 7. ByteDance published the pricing ahead of it, and there is a detail in there that will quietly break your cost model if you carry it over from 2.0.&lt;/p&gt;

&lt;p&gt;Video is quoted per second and metered per token:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;tokens = (input_video_seconds + output_seconds) × width × height × fps / 1024
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;fps is fixed at 24. Multiply by the per-million-token rate and that is the bill.&lt;/p&gt;

&lt;h2&gt;
  
  
  The published rates
&lt;/h2&gt;

&lt;p&gt;USD per million tokens:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;No video input&lt;/th&gt;
&lt;th&gt;With video input&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Seedance 2.5 (480p, 720p)&lt;/td&gt;
&lt;td&gt;10.70&lt;/td&gt;
&lt;td&gt;6.40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Seedance 2.0 (480p, 720p)&lt;/td&gt;
&lt;td&gt;7.00&lt;/td&gt;
&lt;td&gt;4.30&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Seedance 2.0 (1080p)&lt;/td&gt;
&lt;td&gt;7.70&lt;/td&gt;
&lt;td&gt;4.70&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Seedance 2.0 (4K)&lt;/td&gt;
&lt;td&gt;4.00&lt;/td&gt;
&lt;td&gt;2.40&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;2.5 costs 52.9% more per token without video input and 48.8% more with it. Only 480p and 720p are published for 2.5. No 1080p, no 4K, and offline inference reads "not supported yet".&lt;/p&gt;

&lt;p&gt;Look at the 4K row before you move on. It is the cheapest tier per token, 43% below 480p, and it is also the most expensive output on the board, because a 3840×2160 frame carries 19.4 times the pixels of what 480p actually renders. The rate drops 43% while the token count climbs 1940%. Comparing providers by scanning the rate column gets you the wrong answer by roughly a factor of eleven.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 480p frame changed and nobody said so
&lt;/h2&gt;

&lt;p&gt;This is not in any release note. It falls out of dividing ByteDance's own worked examples by their own token rates.&lt;/p&gt;

&lt;p&gt;Their published five-second, 16:9, no-reference examples:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;480p&lt;/th&gt;
&lt;th&gt;720p&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Seedance 2.5&lt;/td&gt;
&lt;td&gt;$0.514 ($0.103/s)&lt;/td&gt;
&lt;td&gt;$1.156 ($0.231/s)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Seedance 2.0&lt;/td&gt;
&lt;td&gt;$0.352 ($0.070/s)&lt;/td&gt;
&lt;td&gt;$0.756 ($0.151/s)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Divide price by token rate to recover the token count, then by 24/1024 to recover pixels:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;tokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;pricePerVideo&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ratePerMillion&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="nx"&gt;e6&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;pixels&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tokens&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nx"&gt;outputSeconds&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;24&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;// Seedance 2.5, 480p:  0.514 / (10.70/1e6) / 5  =  9,607 tokens/sec&lt;/span&gt;
&lt;span class="c1"&gt;//                      9,607 * 1024/24          =  409,899 px  -&amp;gt;  ~854 x 480&lt;/span&gt;
&lt;span class="c1"&gt;// Seedance 2.0, 480p:  0.352 / (7.00/1e6)  / 5  = 10,057 tokens/sec&lt;/span&gt;
&lt;span class="c1"&gt;//                      10,057 * 1024/24         =  429,105 px  -&amp;gt;  ~873 x 491&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;720p resolves to 21,600 tokens per second on both versions, which is exactly 1280 × 720. That the same arithmetic lands on a clean, verifiable number at 720p is what makes the 480p result trustworthy instead of a rounding artifact.&lt;/p&gt;

&lt;p&gt;So 2.5 renders 480p at a true 16:9 854 × 480 while 2.0 uses a slightly taller frame. That 4.5% pixel reduction explains a discrepancy that otherwise looks like a pricing error: &lt;strong&gt;480p rises 46% per second while the token rate rises 53%.&lt;/strong&gt; At 720p, where the frame is unchanged, the two match exactly.&lt;/p&gt;

&lt;p&gt;If you quote customers per second, your 2.0 conversion factor is wrong on 2.5 by about five percent at 480p.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your input clip bills like generated video
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;input_video_seconds&lt;/code&gt; sits inside the same parenthesis as the output. A reference-to-video job pays for the source clip at the same rate as the frames the model invented.&lt;/p&gt;

&lt;p&gt;Providers surface this as a discounted per-second rate for reference mode, which reads like a deal until you total it. Seedance 2.5's published range makes the point on its own: a five-second 720p generation costs $1.244 with a short reference and $4.838 with a 30-second one. Same output, 3.9× the bill.&lt;/p&gt;

&lt;p&gt;2.5 also doubled the input window, from 15 seconds on 2.0 to 30.&lt;/p&gt;

&lt;p&gt;There is a floor too, and the number is unpublished. Their examples price two-second and four-second inputs identically, which implies a minimum around four seconds. The docs point at a spreadsheet calculator rather than stating it. If you resell this, customers sending two-second clips cost you more than your arithmetic predicts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Converting any token rate to per-second
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;TOKENS_PER_SEC&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;480p@2.5&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;854&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;480&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;24&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;//   9,607&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;480p@2.0&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;864&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;496&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;24&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;//  10,044&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;720p&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;    &lt;span class="mi"&gt;1280&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;720&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;24&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;//  21,600&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;1080p&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;  &lt;span class="mi"&gt;1920&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;1080&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;24&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;//  48,600&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;4k&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;     &lt;span class="mi"&gt;3840&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;2160&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;24&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;// 194,400&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;costUsd&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;tier&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;ratePerMillion&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;inputSec&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;outputSec&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;tokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;inputSec&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;outputSec&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;TOKENS_PER_SEC&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;tier&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
  &lt;span class="k"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tokens&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;ratePerMillion&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="nx"&gt;e6&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// Sanity check against the vendor's own example:&lt;/span&gt;
&lt;span class="nf"&gt;costUsd&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;tier&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;720p&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;ratePerMillion&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;7.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;outputSec&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt; &lt;span class="c1"&gt;// 0.756&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last line matches ByteDance's published figure exactly, so the formula is not an approximation.&lt;/p&gt;

&lt;p&gt;One caveat for anyone metering downstream: token counts are estimates until the job finishes. The formula predicted 40,176 for a config where the API returned 40,594, about 1% high. Bill on the returned &lt;code&gt;usage.completion_tokens&lt;/code&gt;, not on your estimate.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I could not determine
&lt;/h2&gt;

&lt;p&gt;Why 1080p and 4K have no published rate for 2.5. Either those tiers do not ship at launch or they arrive separately.&lt;/p&gt;

&lt;p&gt;The exact minimum input duration. It exists and it is not a published number.&lt;/p&gt;

&lt;p&gt;Whether the same formula holds outside the Seedance family. Pixels × duration × fps is a plausible general shape, but the constants and the input-billing rule are not something I would assume for Veo, Kling or Sora without checking. If you have run the same back-solve against those, I would genuinely like to know how it came out.&lt;/p&gt;

&lt;p&gt;Full writeup with the Seedance 2.0 rate card in per-second and per-minute terms: &lt;a href="https://reapi.ai/blog/seedance-2-5-pricing-per-token" rel="noopener noreferrer"&gt;reapi.ai/blog/seedance-2-5-pricing-per-token&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Live parameter reference and a browser playground: &lt;a href="https://reapi.ai/models/seedance-2-0" rel="noopener noreferrer"&gt;reapi.ai/models/seedance-2-0&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>video</category>
      <category>pricing</category>
    </item>
    <item>
      <title>Best Project Management Software for Startups: Match the Tool to How You Work</title>
      <dc:creator>Owais Khan</dc:creator>
      <pubDate>Wed, 05 Aug 2026 03:55:02 +0000</pubDate>
      <link>https://dev.to/reaperoak/best-project-management-software-for-startups-match-the-tool-to-how-you-work-543i</link>
      <guid>https://dev.to/reaperoak/best-project-management-software-for-startups-match-the-tool-to-how-you-work-543i</guid>
      <description>&lt;p&gt;Search "best project management software for startups" and you get the same&lt;br&gt;
dozen names every time: Trello, Asana, ClickUp, Notion, Linear, monday.com,&lt;br&gt;
Basecamp. Ranking them by feature count tells you almost nothing, because they&lt;br&gt;
are not really competing for the same job. The useful question for a startup is&lt;br&gt;
not which tool has the most features. It is two narrower ones: does your work&lt;br&gt;
run through engineering or through the whole company, and does per-seat pricing&lt;br&gt;
or flat-rate pricing fit a headcount that is about to change? Answer those and&lt;br&gt;
the shortlist collapses to two or three.&lt;/p&gt;

&lt;h2&gt;
  
  
  The split that actually decides it
&lt;/h2&gt;

&lt;p&gt;Two forks matter more than any side-by-side feature grid. The first is who the&lt;br&gt;
tool is built for. Issue trackers like Linear are built around the engineering&lt;br&gt;
workflow (issues, cycles, a keyboard-first interface) and feel wrong the moment&lt;br&gt;
a marketer or a founder tries to run a launch plan in them. General work tools&lt;br&gt;
like Asana, ClickUp, monday.com and Trello are built for any team, which makes&lt;br&gt;
them flexible but also less opinionated about how software actually ships.&lt;/p&gt;

&lt;p&gt;The second fork is the shape of the bill. Almost everything in this category&lt;br&gt;
charges per seat per month, so the cost scales directly with hiring. A small&lt;br&gt;
number, Basecamp most notably, offer a flat rate that does not. For a company&lt;br&gt;
planning to double headcount inside a year, that difference can outweigh any&lt;br&gt;
feature comparison.&lt;/p&gt;

&lt;h2&gt;
  
  
  If your team is mostly engineers
&lt;/h2&gt;

&lt;p&gt;For an engineering-led startup, an issue tracker usually beats a general project&lt;br&gt;
tool. Linear's free plan includes unlimited members, two teams and up to 250&lt;br&gt;
issues, which is enough to run a small product team before paying anything; its&lt;br&gt;
Basic plan is $10 per user per month billed yearly and lifts the cap to&lt;br&gt;
unlimited issues and five teams. The trade-off is scope: Linear is deliberately&lt;br&gt;
narrow, so non-engineering work does not fit it well.&lt;/p&gt;

&lt;p&gt;The larger, more familiar alternative is Jira, which startup roundups still name&lt;br&gt;
as the default for software teams despite a steeper learning curve. If most of&lt;br&gt;
your company is not writing code, both are the wrong starting point. You will&lt;br&gt;
spend your first months fighting the tool's assumptions rather than using them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cheapest capable all-rounders
&lt;/h2&gt;

&lt;p&gt;If you want one board-based tool the whole team can adopt, Trello and ClickUp&lt;br&gt;
anchor the low end. Trello's free plan gives you unlimited cards but caps a&lt;br&gt;
Workspace at ten boards and ten collaborators, with 250 automation command runs&lt;br&gt;
a month. Standard is $5 per user per month billed annually and removes the board&lt;br&gt;
cap, while Premium at $10 adds Calendar, Timeline, Table and Dashboard views.&lt;br&gt;
Trello is genuinely easy to adopt, but its weakness is reporting: teams&lt;br&gt;
routinely fall back to spreadsheets for anything above a single board.&lt;/p&gt;

&lt;p&gt;ClickUp aims at the opposite problem by trying to do everything. Its Free&lt;br&gt;
Forever plan is unusual in allowing unlimited members (constrained mainly by&lt;br&gt;
60MB of storage), and its first paid tier, Unlimited, is $7 per user per month&lt;br&gt;
billed yearly. That combination of a free tier that does not punish you for&lt;br&gt;
adding people and a cheap first upgrade is why ClickUp appears on nearly every&lt;br&gt;
budget list. The cost is complexity; the same roundups that recommend it warn&lt;br&gt;
about a learning curve and a busy interface.&lt;/p&gt;

&lt;h2&gt;
  
  
  If your projects already live in your docs
&lt;/h2&gt;

&lt;p&gt;Notion is a documents-and-databases workspace that many startups use for project&lt;br&gt;
tracking rather than a dedicated PM tool. That works well when your projects are&lt;br&gt;
inseparable from your notes, specs and internal wiki. Watch the free plan's&lt;br&gt;
edges: file uploads are capped at 5MB, page history is kept for seven days, and&lt;br&gt;
block limits kick in once you have more than one member. The Plus plan at $10&lt;br&gt;
per member per month lifts those limits.&lt;/p&gt;

&lt;p&gt;Notion is a poor fit if you need Gantt-style timelines or heavy automation out&lt;br&gt;
of the box. It is an excellent fit if your team already lives in Notion docs and&lt;br&gt;
simply wants tasks sitting next to them, rather than in a separate app nobody&lt;br&gt;
opens.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you want structure without inventing it
&lt;/h2&gt;

&lt;p&gt;Asana and monday.com occupy the middle ground: more structure than Trello, less&lt;br&gt;
specialised than Linear. Asana's free Personal plan now caps collaboration at&lt;br&gt;
two users, and its first real team tier, Starter, is $10.99 per user per month&lt;br&gt;
billed annually. That is the tier that unlocks Timeline and Gantt views, forms&lt;br&gt;
and custom fields, so most teams will land there rather than on free.&lt;/p&gt;

&lt;p&gt;monday.com is similar in spirit, with two startup-relevant catches worth knowing&lt;br&gt;
before you fall in love with the demo: paid plans start at a three-seat minimum,&lt;br&gt;
and its free tier is capped at two seats and three boards, which makes the free&lt;br&gt;
plan closer to a trial than a home. Its Standard plan is $12 per seat per month&lt;br&gt;
billed annually, and that is where Timeline, Gantt and guest access appear.&lt;/p&gt;

&lt;h2&gt;
  
  
  If a predictable bill matters more than features
&lt;/h2&gt;

&lt;p&gt;Basecamp is the outlier worth knowing about because of its pricing model, not&lt;br&gt;
its feature list. Its Pro Unlimited plan is a flat $299 per month billed&lt;br&gt;
annually for unlimited users, rather than a per-seat charge. For a team that is&lt;br&gt;
about to keep hiring, a flat rate can be cheaper and far more predictable than&lt;br&gt;
any per-seat tool; the roundups make the same point, that a flat rate starts to&lt;br&gt;
undercut per-seat pricing once headcount climbs. The trade-off is deliberate&lt;br&gt;
simplicity: no Gantt charts, limited automation and little customisation. You&lt;br&gt;
are buying a calm interface and a predictable invoice, not depth.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to check before you commit
&lt;/h2&gt;

&lt;p&gt;Pricing and limits in this category change often, so treat every figure here,&lt;br&gt;
this article included, as something to confirm on the vendor's own pricing page.&lt;br&gt;
Specifically:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Price the tier that has the feature you need&lt;/strong&gt;, not the entry tier. Timeline
and Gantt views, custom fields and real automation sit one tier up on most of
these tools.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Count non-working seats.&lt;/strong&gt; Founders, advisors and contractors who need only
read access move a per-seat total more than the headline number suggests;
check whether guests or viewers are free.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Watch the free-tier cliffs.&lt;/strong&gt; Notion's block limit for teams, Trello's
ten-board cap and Asana's two-user free plan all push you to upgrade sooner
than the marketing implies.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test an export during the trial.&lt;/strong&gt; The tools that are easiest to leave are
the ones you can evaluate honestly in year two, before you have a year of
history locked inside them.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A reasonable default
&lt;/h2&gt;

&lt;p&gt;For a general, non-engineering startup team that wants to spend nothing to start&lt;br&gt;
and little later, ClickUp's unlimited-member free plan and $7 first tier make it&lt;br&gt;
the lowest-risk place to begin, with Trello a simpler alternative if you only&lt;br&gt;
ever need boards. For an engineering-led team, start with Linear and add a&lt;br&gt;
lightweight tool for the rest of the company only when someone actually asks for&lt;br&gt;
one. And if you are about to hire quickly and value a predictable bill over&lt;br&gt;
feature depth, price Basecamp's flat rate against the per-seat tools at your&lt;br&gt;
expected twelve-month headcount before committing. The ranking often flips as&lt;br&gt;
you grow, which is exactly the kind of thing that is cheap to check now and&lt;br&gt;
expensive to discover later.&lt;/p&gt;

</description>
      <category>productivity</category>
      <category>projectmanagement</category>
      <category>saas</category>
      <category>business</category>
    </item>
    <item>
      <title>Test smarter with Snagly: 30 open-source QA skills for AI coding agents</title>
      <dc:creator>Ambreen Khan</dc:creator>
      <pubDate>Wed, 05 Aug 2026 03:54:05 +0000</pubDate>
      <link>https://dev.to/ambytious/test-smarter-with-snagly-30-open-source-qa-skills-for-ai-coding-agents-1571</link>
      <guid>https://dev.to/ambytious/test-smarter-with-snagly-30-open-source-qa-skills-for-ai-coding-agents-1571</guid>
      <description>&lt;p&gt;If you've experimented with AI-driven testing, you've probably lived this cycle: you ask an AI agent to "test the checkout flow," and it does &lt;em&gt;something&lt;/em&gt; — clicks around, declares success, and leaves you unsure what was actually verified. The next day you ask again and it does something different. The browser automation works; the &lt;em&gt;testing discipline&lt;/em&gt; is missing.&lt;/p&gt;

&lt;p&gt;That gap is what &lt;strong&gt;Snagly&lt;/strong&gt; is for.&lt;/p&gt;

&lt;p&gt;Rather than describe it, I pointed it at &lt;a href="https://softwaretestingtrends.com" rel="noopener noreferrer"&gt;softwaretestingtrends.com&lt;/a&gt; — my own production site, nothing fixed beforehand — and recorded the whole thing. It found eleven issues, including a critical accessibility bug on my own signup page. One of its findings turned out to be wrong, and I'll come back to that, because it matters more than the ones it got right.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;📺 &lt;strong&gt;&lt;a href="https://www.youtube.com/watch?v=Ra0v9ois_wU" rel="noopener noreferrer"&gt;Watch the full walkthrough&lt;/a&gt;&lt;/strong&gt; — installed from an empty folder, run against production, ~20 minutes.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;Snagly is a free, MIT-licensed set of 30 skills for AI coding agents — &lt;a href="https://github.com/features/copilot" rel="noopener noreferrer"&gt;GitHub Copilot&lt;/a&gt;, &lt;a href="https://claude.com/claude-code" rel="noopener noreferrer"&gt;Claude Code&lt;/a&gt;, Cursor, Codex and 70+ others — that turn "an AI that can drive a browser" into "an AI that tests like a QA professional." A skill, if you haven't met them yet, is a reusable instruction set that teaches the agent a specific working method — when to use it, what rigor it requires, what evidence to capture, and what it must never do.&lt;/p&gt;

&lt;p&gt;Each skill in Snagly has one job, and they hand off to each other the way a real testing practice does:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;start-testing&lt;/code&gt;&lt;/strong&gt; is the front door — say "what can you test here?" and it routes you to the right skill, checking prerequisites before handing off.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Discovery &amp;amp; strategy&lt;/strong&gt;: &lt;code&gt;scenario-mapper&lt;/code&gt; explores your site and produces a prioritized list of test scenarios; &lt;code&gt;test-case-writer&lt;/code&gt; expands any of them into a reviewable spec; &lt;code&gt;test-plan&lt;/code&gt; sets strategy, cadence, and release exit criteria; &lt;code&gt;qa-onboarding&lt;/code&gt; writes the guide for your next hire.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Execution&lt;/strong&gt;: &lt;code&gt;flow-runner&lt;/code&gt; drives real user journeys step by step, asserting outcomes (not just that clicks happened) and capturing evidence the moment anything fails. &lt;code&gt;crud-tester&lt;/code&gt; is the only skill allowed to mutate data — under strict rules we'll get to. &lt;code&gt;e2e-codegen&lt;/code&gt; converts a &lt;em&gt;verified&lt;/em&gt; flow into a permanent &lt;code&gt;@playwright/test&lt;/code&gt; spec.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The defect loop&lt;/strong&gt;: &lt;code&gt;bug-triage&lt;/code&gt; reproduces a suspected bug and establishes a minimal repro with an evidence bundle; &lt;code&gt;bug-creator&lt;/code&gt; files it in Jira — deduplicated against existing tickets first; &lt;code&gt;fix-verifier&lt;/code&gt; re-runs the repro on later builds and tells you FIXED, STILL BROKEN, or REGRESSED. &lt;code&gt;bug-analyzer&lt;/code&gt; works the other direction: an existing ticket in, ranked root-cause hypotheses out.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Specialized checks&lt;/strong&gt;: &lt;code&gt;network-assertion&lt;/code&gt; (mock API failures and assert on real traffic), &lt;code&gt;cross-browser-matrix&lt;/code&gt;, &lt;code&gt;auth-session-audit&lt;/code&gt;, &lt;code&gt;form-fuzzing&lt;/code&gt;, &lt;code&gt;email-verification&lt;/code&gt; — each executing one kind of check properly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Site-wide audits&lt;/strong&gt;: &lt;code&gt;accessibility-audit&lt;/code&gt; (axe-core plus the manual checks axe can't do), &lt;code&gt;performance-audit&lt;/code&gt; (Core Web Vitals), &lt;code&gt;seo-audit&lt;/code&gt;, &lt;code&gt;i18n-audit&lt;/code&gt;, &lt;code&gt;link-audit&lt;/code&gt;, and &lt;code&gt;security-hygiene&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Visual &amp;amp; design&lt;/strong&gt;: &lt;code&gt;visual-snapshot&lt;/code&gt; captures a reviewable gallery of every page; &lt;code&gt;visual-regression&lt;/code&gt; pixel-diffs two captures; &lt;code&gt;figma-compare&lt;/code&gt; checks the built UI against its Figma design, field by field.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Synthesis&lt;/strong&gt;: &lt;code&gt;report-generator&lt;/code&gt; turns everything the other skills produced — across sessions, across a whole testing cycle — into one prioritized report you can send to your team.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Under the hood, the browser work runs on Playwright (via the Playwright MCP server or &lt;code&gt;@playwright/cli&lt;/code&gt;), the design side uses the Figma MCP server, and the Jira family talks to Jira Cloud through a dependency-free Python client.&lt;/p&gt;

&lt;h2&gt;
  
  
  The opinions baked in
&lt;/h2&gt;

&lt;p&gt;Tools are easy; discipline is hard. The value of these skills is less in what they &lt;em&gt;do&lt;/em&gt; and more in what they &lt;em&gt;refuse to do&lt;/em&gt;. A few principles run through the whole toolkit:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Evidence over vibes.&lt;/strong&gt; Every finding cites what was actually observed — the screenshot, the console error, the network response. A bug isn't a bug until it has a minimal repro and a reproducibility count. &lt;code&gt;bug-creator&lt;/code&gt; will actively route an unverified finding back through &lt;code&gt;bug-triage&lt;/code&gt; before filing, because one withdrawn false positive costs more credibility than ten good tickets earn.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verified and inferred are never confused.&lt;/strong&gt; A mocked API response, a lab performance number, and a test case nobody has executed yet are all clearly labeled as such. &lt;code&gt;e2e-codegen&lt;/code&gt; refuses to generate test code from a scenario that's never actually been run — that would bake untested assumptions into something that looks authoritative.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mutations are contained.&lt;/strong&gt; Only one skill is allowed to create, edit, or delete data, and only in a tenant you've explicitly named as safe. Every record it creates carries a run marker, and cleanup is itself a test. Every Jira write in the toolkit is dry-run by default — nothing is filed, commented, or transitioned until you've seen the exact payload and said yes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Explicit about what wasn't covered.&lt;/strong&gt; Every run report states its scope and its gaps rather than implying completeness. And two things are deliberately out of scope: aesthetic judgment (not checkable the way everything else is) and anything resembling penetration testing — the fuzzing and hygiene skills draw a hard line at injection payloads and exploit attempts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Knowledge compounds.&lt;/strong&gt; A target profile (&lt;code&gt;targets/*.yaml&lt;/code&gt;) records everything the toolkit learns about your app — the login quirks, the safe tenant, the known console noise, the field that rejects "+" in phone numbers — so every run makes the next run cheaper instead of rediscovering the same facts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Getting started
&lt;/h2&gt;

&lt;p&gt;You'll need Node.js, an agent that supports skills, and the Playwright CLI (&lt;code&gt;npm install -g @playwright/cli@latest&lt;/code&gt;, then &lt;code&gt;playwright-cli install --skills&lt;/code&gt;). Then install Snagly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add softwaretestingtrends/snagly &lt;span class="nt"&gt;--all&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On Claude Code you can install it as a plugin instead, which handles updates for you:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;/plugin marketplace add softwaretestingtrends/snagly
/plugin &lt;span class="nb"&gt;install &lt;/span&gt;snagly@snagly
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Point it at your app by copying &lt;code&gt;targets/example.yaml&lt;/code&gt; into your project and filling in the base URL, where credentials live (env vars — never in the file), and the login recipe. Then just ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What can you test here?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;code&gt;start-testing&lt;/code&gt; takes it from there — you don't have to know which of the thirty skills you need.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it found on my own site
&lt;/h2&gt;

&lt;p&gt;Here's the part I can't fake. My own site, in production, nothing fixed beforehand:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A sitemap that doesn't exist.&lt;/strong&gt; My &lt;code&gt;robots.txt&lt;/code&gt; has been pointing search engines at a 404 for months.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A critical accessibility bug on the signup page.&lt;/strong&gt; The show/hide password toggles had no accessible name, so screen-reader users couldn't tell what they did — on the page where people create accounts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nine pages sharing one title tag&lt;/strong&gt; and one meta description, no canonicals, no Open Graph tags. That last one is why links to my site had been rendering blank previews.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A pricing bug nobody would find by hand.&lt;/strong&gt; One course is currently free. Shopify charges zero, correctly. The button on my own site still said $24.99 — caught only because it compared what the button &lt;em&gt;claims&lt;/em&gt; against what checkout actually &lt;em&gt;does&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A background 406 on every enrolled course page&lt;/strong&gt;, invisible in the UI, failing on every single load.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And a result I didn't expect: performance came back &lt;strong&gt;clean&lt;/strong&gt;. Core Web Vitals good across the board, with an unprompted note that these were lab numbers, not the field data Google grades against. A tool that only ever finds problems isn't measuring anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  The finding that was wrong
&lt;/h2&gt;

&lt;p&gt;One report said the login and signup pages had no visible focus indicator — a real WCAG failure if true. I tabbed through, and I could plainly see an orange focus ring.&lt;/p&gt;

&lt;p&gt;It was a false positive, and the cause is instructive. The check read computed styles after focusing elements &lt;em&gt;programmatically&lt;/em&gt;, which doesn't reliably trigger &lt;code&gt;:focus-visible&lt;/code&gt;. Frameworks that build their focus ring from CSS variables — Tailwind, in my case — read as fully transparent in that state while actually rendering perfectly. The tool measured a ring that hadn't been asked to appear yet.&lt;/p&gt;

&lt;p&gt;I shipped a fix the same day: that check may no longer report a failure from computed styles at all. It has to press Tab for real and compare a focused screenshot against an unfocused one.&lt;/p&gt;

&lt;p&gt;I'm telling you this because it's the honest answer to the question you should be asking about any AI testing tool: &lt;em&gt;what happens when it's wrong?&lt;/em&gt; Here, a human caught it in thirty seconds, and the toolkit got permanently better. Nineteen of Snagly's improvements so far came from exactly this — using it for real and fixing what broke.&lt;/p&gt;

&lt;p&gt;The full prerequisites (including the optional Figma, Jira, and email-testing setups) are in the &lt;a href="https://github.com/softwaretestingtrends/snagly" rel="noopener noreferrer"&gt;README&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the name
&lt;/h2&gt;

&lt;p&gt;If you've worked a release the traditional way, you know the &lt;strong&gt;snag list&lt;/strong&gt; — the running register of every defect, rough edge, and "that's not quite right" that stands between a build and a sign-off. Snagly is a toolkit for building that list &lt;em&gt;properly&lt;/em&gt;: every snag caught with evidence, reproduced before it's reported, tracked until it's verified fixed. (And yes, it's still in service of the Software Testing Trends motto — &lt;em&gt;learn smarter, test better&lt;/em&gt;.)&lt;/p&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;Snagly is MIT-licensed and open to contributions — new skills, better target-profile patterns, gotchas from your own Jira or app under test. &lt;a href="https://github.com/softwaretestingtrends/snagly" rel="noopener noreferrer"&gt;Star the repo&lt;/a&gt; to follow releases (each one is tagged and versioned; updating is one &lt;code&gt;/plugin marketplace update snagly&lt;/code&gt; away). It's also been submitted to the Claude Code community plugin marketplace — once listed there, it'll be browsable directly from &lt;code&gt;/plugin&lt;/code&gt; with no marketplace-add step.&lt;/p&gt;

&lt;p&gt;Part two is coming: I fix the issues above, then put Snagly back on the site to verify the fixes independently and close the Jira tickets it opened.&lt;/p&gt;

&lt;p&gt;I'm also putting together a deeper guide on running an AI-augmented QA practice — how to adapt these skills to your own app, build new ones, and roll this out to a team. If you want that when it's ready, &lt;a href="https://www.youtube.com/@softwaretestingtrends" rel="noopener noreferrer"&gt;subscribe here&lt;/a&gt; — and if you take Snagly for a spin this week, I'd genuinely love to hear what broke, what surprised you, and what's missing.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;— Ambreen Khan, &lt;a href="https://softwaretestingtrends.com" rel="noopener noreferrer"&gt;Software Testing Trends&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>testing</category>
      <category>ai</category>
      <category>playwright</category>
      <category>agentskills</category>
    </item>
    <item>
      <title>Episode 6 — Watching Something You Can't See</title>
      <dc:creator>surajrkhonde</dc:creator>
      <pubDate>Wed, 05 Aug 2026 03:40:11 +0000</pubDate>
      <link>https://dev.to/surajrkhonde/episode-6-watching-something-you-cant-see-15eb</link>
      <guid>https://dev.to/surajrkhonde/episode-6-watching-something-you-cant-see-15eb</guid>
      <description>&lt;p&gt;&lt;em&gt;Week 3. "The deploy is done. Everything's green. Now what am I actually supposed to be looking at?"&lt;/em&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  Previously
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Runner
    ↓
Cache
    ↓
Artifact
    ↓
Deployment

Today
    ↓
Monitoring
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; The canary rolled out fine yesterday. 100% traffic, all healthy. I closed my laptop. Was that wrong?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; Not wrong, exactly. But let me ask you something first. Your service is running on a server somewhere. Right now, this second — is it healthy?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; I mean... I assume so? Nobody's messaged me.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; "Nobody's messaged me" isn't an answer. It's the absence of one. That's the entire problem monitoring exists to solve.&lt;/p&gt;




&lt;h3&gt;
  
  
  The Thing Nobody Says Out Loud
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; Here's an uncomfortable fact about production systems: you cannot see them. Not directly. You're not standing next to the server, watching electricity move through it. Everything you know about whether it's healthy is a claim — something a piece of software told you, that you're choosing to trust.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; That sounds obvious when you say it, but I don't think I've ever actually thought about it that way.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; Most engineers don't, until the gap between "the system told me it's fine" and "the system is actually fine" bites them. Monitoring is the discipline of shrinking that gap — of making sure what you're told is close to what's actually true, and told to you fast enough to matter.&lt;/p&gt;




&lt;h3&gt;
  
  
  📒 Senior Engineer's Notebook
&lt;/h3&gt;

&lt;p&gt;&lt;em&gt;You don't monitor a system because you don't trust it. You monitor it because you can't see it. Trust isn't the issue — visibility is.&lt;/em&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  The Car Dashboard Analogy
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; Can you make this concrete?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; Think about driving a car. You can't see the engine. You can't see the oil level, the coolant temperature, how much fuel is actually left in the tank, mid-drive. All of that is invisible to you, sealed inside metal, while you're doing 100 km/h.&lt;/p&gt;

&lt;p&gt;So the car gives you a dashboard. Speed, fuel, engine temperature, warning lights. You're not watching the engine. You're watching &lt;em&gt;proxies&lt;/em&gt; for the engine — numbers and lights standing in for things you can't directly observe.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; And a production system is the engine. Monitoring is the dashboard.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; Exactly that. And just like a car, the dangerous failure mode isn't "the dashboard shows a problem." It's "the dashboard shows everything's fine, and it's wrong." A dashboard that lies is worse than no dashboard — because no dashboard, at least, you know you're flying blind.&lt;/p&gt;




&lt;h3&gt;
  
  
  What Actually Gets Measured
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; Okay, so — what goes on the dashboard? What do you actually watch?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; There's a well-known starting set, sometimes called the four golden signals. Not the only things worth tracking, but the ones that catch the most real problems the fastest.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Latency&lt;/strong&gt; — how long requests take. Not just the average — the average can look perfectly healthy while 5% of your users wait eight seconds. You want percentiles: p50, p95, p99. What's typical, and what's the bad end of typical.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Traffic&lt;/strong&gt; — how many requests you're actually getting. Without this number, every other number is meaningless. An error rate of 2% means something completely different at ten requests a minute versus ten thousand.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Errors&lt;/strong&gt; — the rate of requests failing. Not just 500s — a request that "succeeds" with the wrong data is a failure your server doesn't know to report as one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Saturation&lt;/strong&gt; — how close your system is to its limit. CPU, memory, database connections, queue depth. A system can have zero errors right now and still be forty seconds from falling over.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; That last one feels different from the others. The others are about what's happening. Saturation is about what's &lt;em&gt;about&lt;/em&gt; to happen.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; That's a sharp distinction, and it's exactly right. Latency, traffic, and errors tell you the present. Saturation is close to the only one of the four that gives you a warning &lt;em&gt;before&lt;/em&gt; the present turns bad.&lt;/p&gt;




&lt;h3&gt;
  
  
  🪞 If I asked you this in an interview
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;"What are the four golden signals, and why those four specifically?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Latency, traffic, errors, and saturation. Together they answer the questions that matter most during an incident: is it slow, how much load is it under, is it actually failing, and how close is it to running out of capacity. Most production problems show up in at least one of these four before they show up anywhere else — which is why they're the default starting point rather than an exhaustive list.&lt;/em&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  Monitoring vs Alerting
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; Is a dashboard the same thing as monitoring, then? Just — numbers on a screen?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; A dashboard is monitoring's &lt;em&gt;visible&lt;/em&gt; half. But a dashboard only helps if someone's staring at it at the exact moment something breaks. Nobody's staring at a dashboard at 3 AM.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; So how does anyone find out at 3 AM?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; That's &lt;strong&gt;alerting&lt;/strong&gt; — the other half. You define a threshold: error rate above 5% for two minutes, latency p99 above 3 seconds, saturation above 90%. Cross that threshold, and instead of waiting for a human to notice a graph, the system pages someone. This is where PagerDuty from Episode 1 actually connects to everything we've built since — the pager doesn't go off because a human is watching. It goes off because a machine was, continuously, and a human wasn't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; So monitoring is the measuring. Alerting is the "go wake somebody up" part.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; Precisely. Monitoring without alerting is a dashboard nobody's watching. Alerting without monitoring doesn't exist — there's nothing to alert on.&lt;/p&gt;

&lt;p&gt;And dashboards aren't only useful during an outage, worth saying explicitly. A healthy-looking dashboard, watched over days and weeks, is how teams catch trends before they become incidents at all — latency creeping up 5% every release, memory usage climbing slightly with each deploy, error rate drifting from 0.1% to 0.4% over a month with nothing dramatic enough to trip an alert. Nothing there pages anyone. But someone glancing at trends, not just thresholds, catches the slow version of the same story the fast version tells at 3 AM.&lt;/p&gt;




&lt;h3&gt;
  
  
  📝 Production Note
&lt;/h3&gt;

&lt;p&gt;Remember correlation IDs and structured logs, from Episode 1? That's not a separate topic from monitoring — it's the raw material monitoring is built from. A dashboard showing "errors spiked at 2:14 PM" tells you &lt;em&gt;that&lt;/em&gt; something broke. Structured logs, searchable by correlation ID, are what let you find out &lt;em&gt;which&lt;/em&gt; request, for &lt;em&gt;which&lt;/em&gt; user, touching &lt;em&gt;which&lt;/em&gt; service — the difference between knowing something's wrong and knowing what to actually fix.&lt;/p&gt;

&lt;p&gt;In modern observability, engineers usually lean on three sources of truth together, not one alone: &lt;strong&gt;metrics&lt;/strong&gt; tell you &lt;em&gt;something&lt;/em&gt; is wrong — a number crossed a line. &lt;strong&gt;Logs&lt;/strong&gt; help explain &lt;em&gt;what happened&lt;/em&gt; — the specific events around it. &lt;strong&gt;Traces&lt;/strong&gt; show &lt;em&gt;where the request spent its time&lt;/em&gt; as it moved across multiple services. You'll meet traces properly once we're deep into distributed systems — for now, just know the three aren't competing tools, they're three different angles on the same question.&lt;/p&gt;




&lt;h3&gt;
  
  
  The Alert Fatigue Trap
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; This seems easy to over-do, though. Why not just alert on everything, so nothing slips through?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; Because a human being can only take so many 3 AM pages before they stop trusting them. That's called &lt;strong&gt;alert fatigue&lt;/strong&gt;, and it's one of the most common ways monitoring quietly fails — not by missing a real incident, but by burying it under noise.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; How does that actually play out?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; Say your team gets paged fifteen times a week. Twelve of those are noise — a brief blip that self-resolved, a threshold set too aggressively, a known flaky check. After a few weeks of that, the on-call engineer starts silencing pages before fully reading them. Reflex, not laziness — self-preservation.&lt;/p&gt;

&lt;p&gt;Then, one night, page sixteen is the real one. And it gets the same half-second of attention as the twelve before it that didn't matter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; So too many alerts is almost worse than too few.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; In a very specific sense, yes — because too few means a known gap. Too many means a false sense of coverage while the actual signal quietly drowns.&lt;/p&gt;

&lt;p&gt;Two terms are worth having exact language for here, because you'll hear them constantly. A &lt;strong&gt;false positive&lt;/strong&gt; is an alert that fires when nothing's actually wrong — that's the noise causing fatigue. A &lt;strong&gt;false negative&lt;/strong&gt; is the opposite and more dangerous failure: something &lt;em&gt;is&lt;/em&gt; wrong, and no alert fires at all. Tuning alerts is really just a constant tradeoff between the two — tighten thresholds to catch more real problems, and you risk more false positives; loosen them to reduce noise, and you risk missing something real. There's no setting that eliminates both.&lt;/p&gt;




&lt;h3&gt;
  
  
  👦 Junior Framing / 👨‍🦳 Senior Framing
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Junior framing:&lt;/strong&gt; "More alerts means better monitoring."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior framing:&lt;/strong&gt; More alerts means more noise, unless each one is something a human genuinely needs to act on right now. A good alert is a promise: if this fires, it matters, and someone needs to do something about it immediately. Break that promise enough times and the alert stops meaning anything.&lt;/p&gt;




&lt;h3&gt;
  
  
  SLOs and Error Budgets
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; How do teams decide what "healthy" even means numerically? Is 99% uptime good? 99.9%?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; This is where &lt;strong&gt;SLOs&lt;/strong&gt; — service level objectives — come in. A team explicitly decides: "99.9% of requests should succeed, measured over 30 days." That's not aspirational marketing language. It's a number the team commits to, and measures against.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; Why not just aim for 100%?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; Because 100% is enormously expensive, and past a certain point, chasing it stops making users happier and starts just slowing the team down. 99.9% still means about 43 minutes of downtime a month are considered acceptable — not desired, but acceptable, a deliberately budgeted amount of failure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; Budgeted failure. That's a strange phrase.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; It's called an &lt;strong&gt;error budget&lt;/strong&gt;, and the phrase is supposed to feel strange — it's meant to change behavior. If you've used up this month's error budget on a risky deploy that caused an outage, that's a signal: slow down, stabilize, stop taking risky bets for a while. If you've got budget left, you have room to ship faster and take more chances. It turns "is it safe to deploy" from a gut feeling into an actual number.&lt;/p&gt;




&lt;h3&gt;
  
  
  🚨 Beginner Mistakes
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;"I'll alert on CPU usage being high."&lt;/strong&gt; High CPU isn't inherently bad — sometimes it means the system is efficiently using the resources it has. Alert on the thing that actually matters to users: is latency degrading, are requests failing. CPU is a diagnostic detail, not the actual problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"I'll set every alert threshold as tight as possible, to catch things early."&lt;/strong&gt; This is how alert fatigue starts. A threshold that fires on normal, healthy variation trains people to ignore it — right up until the day it fires for something real.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"The dashboard looks fine, so we're done watching."&lt;/strong&gt; A dashboard is a snapshot of now. The four golden signals can look perfectly healthy one minute before a slow memory leak crosses a threshold. Monitoring is a continuous practice, not a one-time check after deploying.&lt;/p&gt;




&lt;h3&gt;
  
  
  🏢 Office Reality
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;"We're flying blind on this service."&lt;/strong&gt; No meaningful monitoring exists for it — nobody would know if it broke until a user complained.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"That alert is noisy."&lt;/strong&gt; It fires often enough, for things that don't actually matter, that people have started ignoring it. A polite way of saying "this alert has stopped doing its job."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"We burned our error budget."&lt;/strong&gt; The team exceeded their allowed failure rate for the period — often the trigger for pausing risky launches until things stabilize.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Check the dashboards before you page anyone."&lt;/strong&gt; Standard on-call instinct — confirm what's actually happening before waking up a second person for something that might resolve on its own.&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  The Alert That Almost Wasn't
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; Has an alert actually caught something real for you — not in theory, an actual time?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; More than once, but here's a clean one. A saturation alert — database connection pool climbing toward its limit — fired quietly on a Tuesday afternoon. No errors yet. No user complaints. Latency barely moved.&lt;/p&gt;

&lt;p&gt;Easy to ignore. Nothing was actually broken yet.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; But someone looked anyway.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; Someone looked anyway, because the alert existed and the team trusted it enough not to reflexively dismiss it — which, notice, is the entire alert fatigue conversation paying off in real time. Turned out a recent deploy had a connection leak — every request opened a database connection and a rare error path failed to close it. Slow leak. Would've taken maybe three more hours to actually exhaust the pool and start failing every single request, in the middle of the evening peak.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; So the alert fired before anything was actually broken.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; That's the entire point of saturation as a signal. It's not telling you what's wrong right now. It's telling you what's about to be wrong, while there's still time to do something other than panic.&lt;/p&gt;




&lt;h3&gt;
  
  
  🎯 Interview Perspective
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Interviewer:&lt;/strong&gt; How would you decide what to put an alert on?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Weak answer:&lt;/strong&gt; Alert on anything that could go wrong.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Strong answer:&lt;/strong&gt; Alert on symptoms that directly affect users or clearly predict an imminent failure — elevated error rate, degraded latency, saturation approaching a limit — not on every internal metric that could theoretically be interesting. Every alert should represent something a human genuinely needs to act on right now; anything less specific belongs on a dashboard for investigation, not in someone's pocket at 3 AM.&lt;/p&gt;




&lt;h3&gt;
  
  
  🎤 Explain It In One Minute
&lt;/h3&gt;

&lt;p&gt;Imagine explaining this to a teammate — without using the words &lt;em&gt;monitoring&lt;/em&gt;, &lt;em&gt;alerting&lt;/em&gt;, or &lt;em&gt;SLO&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Why is "the deploy went fine" not the same claim as "the system is healthy"?&lt;/p&gt;

&lt;p&gt;If you can answer that without reaching for the vocabulary, you understand the idea driving this entire episode.&lt;/p&gt;




&lt;h3&gt;
  
  
  Whiteboard Moment
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Production system running (invisible to you directly)
    ↓
Metrics collected continuously
    (latency, traffic, errors, saturation)
    ↓
Dashboards — visible, but nobody's always watching
    ↓
Thresholds defined against SLOs
    ↓
Threshold crossed → alert fires → someone paged
    ↓
Structured logs + correlation IDs → find the specific cause
    ↓
Fixed, or escalated toward an incident
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; So this whole episode is really just: you can't see it, so you measure it, and you don't wait for a human to notice the measurement.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; That's the entire discipline in one sentence. Everything else — golden signals, SLOs, error budgets — is just making that sentence precise enough to actually act on.&lt;/p&gt;




&lt;h3&gt;
  
  
  What You Should Be Able to Explain Now
&lt;/h3&gt;

&lt;p&gt;&lt;em&gt;(Without looking at Google)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Can you explain:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why "the deploy succeeded" and "the system is healthy" are different claims?&lt;/li&gt;
&lt;li&gt;What the four golden signals are, and why saturation is different from the other three?&lt;/li&gt;
&lt;li&gt;The difference between monitoring and alerting?&lt;/li&gt;
&lt;li&gt;Why too many alerts can be more dangerous than too few?&lt;/li&gt;
&lt;li&gt;What an SLO and an error budget actually are, in plain terms?&lt;/li&gt;
&lt;li&gt;Why saturation alerts can fire before anything is technically broken yet — and why that's the point?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If yes — you understand production visibility the way it's actually practiced, not just as a dashboard you glance at after a deploy.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; Okay, so say an alert &lt;em&gt;does&lt;/em&gt; fire. For real, not a false alarm. What happens next?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; Now we're back exactly where Episode 1 started — Friday, 7:02 PM, PagerDuty going off. Except this time, you're not watching from the outside. You already know what a webhook is, what a runner is, what an artifact is, how it got deployed, and what the dashboard was supposed to be telling someone the whole time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Junior Engineer:&lt;/strong&gt; So next time, I'm not just shadowing. I'm actually part of figuring it out.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Engineer:&lt;/strong&gt; That's the idea. Incident response, properly this time — from the inside.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(End of Episode 6)&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cicd</category>
      <category>devops</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>LLM Latency Budget: Make AI Features Feel Fast Without Burning Money</title>
      <dc:creator>Jack M</dc:creator>
      <pubDate>Wed, 05 Aug 2026 03:38:03 +0000</pubDate>
      <link>https://dev.to/jackm-singularity/llm-latency-budget-make-ai-features-feel-fast-without-burning-money-3mc3</link>
      <guid>https://dev.to/jackm-singularity/llm-latency-budget-make-ai-features-feel-fast-without-burning-money-3mc3</guid>
      <description>&lt;p&gt;A slow AI feature does not feel smart. It feels broken.&lt;/p&gt;

&lt;p&gt;That is the uncomfortable truth many AI SaaS builders hit after the demo works. The prototype answers well, the agent can call tools, and the RAG pipeline looks impressive. Then real users arrive. Prompts get longer. Queues form. Streaming starts late. One tenant uploads huge documents. Another runs bulk jobs at noon. Suddenly the same workflow that felt magical in testing feels like a spinner with an invoice attached.&lt;/p&gt;

&lt;p&gt;The fix is not simply “use a faster model.” You need an &lt;strong&gt;LLM latency budget&lt;/strong&gt;: a small set of rules that says how fast each AI workflow must feel, how many tokens it can spend, when to stream, when to cache, when to route to another model, and when to stop before cost and latency drift together.&lt;/p&gt;

&lt;p&gt;This guide is for solo SaaS developers, micro SaaS builders, and AI SaaS teams shipping production features with LLM APIs, RAG, agents, or self-hosted models.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why latency budgets matter now
&lt;/h2&gt;

&lt;p&gt;AI platform news points in the same direction: builders are moving from chat demos to production workflows. Agent tools, web context APIs, voice agents, coding assistants, and RAG platforms are all getting more capable. At the same time, inference cost and reliability are under pressure.&lt;/p&gt;

&lt;p&gt;Latency is now a product metric. Inference efficiency is becoming a business metric. Yet many articles stop at TTFT, TPOT, quantization, batching, or model serving. Fewer show how a SaaS builder turns those ideas into a product-level budget with code, dashboards, fallbacks, and customer-safe limits.&lt;/p&gt;

&lt;h2&gt;
  
  
  The simple model: TTFT, TPOT, and total time
&lt;/h2&gt;

&lt;p&gt;You do not need a PhD in serving systems to start. Track three numbers.&lt;/p&gt;

&lt;h3&gt;
  
  
  Time to First Token
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Time to First Token (TTFT)&lt;/strong&gt; is the delay between the user action and the first streamed token. It includes network time, queue time, provider overhead, tool setup, retrieval, and the model’s prefill phase.&lt;/p&gt;

&lt;p&gt;High TTFT is why a chat box feels dead.&lt;/p&gt;

&lt;h3&gt;
  
  
  Time Per Output Token
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Time Per Output Token (TPOT)&lt;/strong&gt; is the average time between generated tokens after the first token appears.&lt;/p&gt;

&lt;p&gt;High TPOT is why streaming feels like a dripping tap.&lt;/p&gt;

&lt;h3&gt;
  
  
  End-to-end latency
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;End-to-end latency&lt;/strong&gt; is the full time from request to final answer.&lt;/p&gt;

&lt;p&gt;A rough formula is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;end_to_end_latency = TTFT + (output_tokens - 1) * TPOT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That formula is not perfect for every provider, but it is good enough to reason about the user experience.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build budgets by workflow, not by model
&lt;/h2&gt;

&lt;p&gt;A common mistake is to set one global target like “AI responses must finish in 5 seconds.” That sounds clean but fails fast.&lt;/p&gt;

&lt;p&gt;Different workflows need different budgets.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workflow&lt;/th&gt;
&lt;th&gt;User expectation&lt;/th&gt;
&lt;th&gt;Suggested latency budget&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Inline autocomplete&lt;/td&gt;
&lt;td&gt;Feels instant&lt;/td&gt;
&lt;td&gt;TTFT under 300ms, very short output&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chat answer&lt;/td&gt;
&lt;td&gt;Starts quickly&lt;/td&gt;
&lt;td&gt;TTFT under 1.5s, stream response&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RAG answer with citations&lt;/td&gt;
&lt;td&gt;Trust matters&lt;/td&gt;
&lt;td&gt;TTFT under 3s, final answer under 15s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent with tool calls&lt;/td&gt;
&lt;td&gt;Progress matters&lt;/td&gt;
&lt;td&gt;First status under 1s, step updates every few seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bulk document task&lt;/td&gt;
&lt;td&gt;Completion matters&lt;/td&gt;
&lt;td&gt;Async job, no chat-style waiting&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The key is to budget for the &lt;strong&gt;experience&lt;/strong&gt;, not the raw model call.&lt;/p&gt;

&lt;p&gt;A user can forgive a 40-second background report if the UI says what is happening. The same user may abandon a 6-second inline writing assistant if nothing appears.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical LLM latency budget template
&lt;/h2&gt;

&lt;p&gt;Create a budget object for each AI workflow.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"workflow"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"support_rag_answer"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"max_ttft_ms"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2500&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"max_total_ms"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;15000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"max_input_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;12000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"max_output_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;900&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"stream"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cache_policy"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"semantic_and_exact"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"fallback_model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"fast_general_model"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"requires_citations"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"async_after_ms"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;12000&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This turns “make it faster” into engineering constraints. Your app can now decide whether to trim context, stream, route to a faster model, switch to async, reject an oversized request, or use a cached answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Instrument every request
&lt;/h2&gt;

&lt;p&gt;Start by logging latency and token data for every AI request. Do this before buying another tool or changing providers.&lt;/p&gt;

&lt;p&gt;Here is a small TypeScript-style example.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;LlmTrace&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;requestId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;inputTokens&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;outputTokens&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;ttftMs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;totalMs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;costUsd&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;cacheHit&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;boolean&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;success&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;timeout&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;error&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;runWithTrace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;started&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;firstTokenAt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;output&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;stream&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;llm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;fast-general&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;max_tokens&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;700&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="k"&gt;await &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;chunk&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;firstTokenAt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;firstTokenAt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="nx"&gt;output&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="nf"&gt;sendToClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;finished&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;trace&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;LlmTrace&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;requestId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;crypto&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;randomUUID&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="na"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;fast-general&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;inputTokens&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;estimateTokens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;outputTokens&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;estimateTokens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;output&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;ttftMs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;firstTokenAt&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;firstTokenAt&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;started&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;totalMs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;finished&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;started&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;costUsd&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;estimateCost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;output&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;cacheHit&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;success&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;

  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;saveTrace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;trace&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;output&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep the trace simple. If you capture request ID, tenant ID, workflow, model, tokens, TTFT, total time, cost, cache hit, and status, you can answer most early performance questions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Control input tokens before touching infrastructure
&lt;/h2&gt;

&lt;p&gt;Long prompts hurt TTFT. Long context means more work before the first token appears.&lt;/p&gt;

&lt;p&gt;For AI SaaS products, input bloat usually comes from full chat history, too many RAG chunks, raw HTML, unused tool descriptions, repeated system instructions, or entire customer records when only a few fields matter. Before optimizing GPUs or switching vendors, cut useless context.&lt;/p&gt;

&lt;p&gt;Use a context packer.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;ContextItem&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;priority&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;tokenEstimate&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;packContext&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;items&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ContextItem&lt;/span&gt;&lt;span class="p"&gt;[],&lt;/span&gt; &lt;span class="nx"&gt;maxTokens&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;sorted&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[...&lt;/span&gt;&lt;span class="nx"&gt;items&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;sort&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;priority&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;priority&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;selected&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ContextItem&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[];&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;used&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;item&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;used&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tokenEstimate&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;maxTokens&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;continue&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="nx"&gt;selected&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="nx"&gt;used&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tokenEstimate&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;selected&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is not fancy. That is the point. A basic priority-based packer often beats “send everything and hope.”&lt;/p&gt;

&lt;p&gt;For RAG, use fewer, better chunks. For agents, expose fewer tools per step. For browser automation, clean the page before putting it into the prompt.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cap output tokens by job type
&lt;/h2&gt;

&lt;p&gt;Output tokens drive total latency and cost. Many AI features do not need long answers.&lt;/p&gt;

&lt;p&gt;Set output caps by workflow:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Rewrite suggestion: 120 tokens&lt;/li&gt;
&lt;li&gt;Error explanation: 250 tokens&lt;/li&gt;
&lt;li&gt;Support answer: 700 tokens&lt;/li&gt;
&lt;li&gt;Technical plan: 1,200 tokens&lt;/li&gt;
&lt;li&gt;Background report: async job with a larger cap&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Also give the model a structure that discourages rambling.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Answer in this format:
1. Direct answer: 2 sentences max
2. Steps: up to 5 bullets
3. Caveat: 1 short note if needed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This improves scannability and reduces token drift.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use streaming for perception, not as a bandage
&lt;/h2&gt;

&lt;p&gt;Streaming can make an AI feature feel faster, but it does not fix everything.&lt;/p&gt;

&lt;p&gt;Use streaming when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The user is reading generated text&lt;/li&gt;
&lt;li&gt;The answer may take more than 2 seconds&lt;/li&gt;
&lt;li&gt;Partial output is useful&lt;/li&gt;
&lt;li&gt;You can show citations or tool results after the draft begins&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Do not rely on streaming when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The workflow must return valid JSON&lt;/li&gt;
&lt;li&gt;The user needs a single deterministic result&lt;/li&gt;
&lt;li&gt;The model must complete tool calls before saying anything&lt;/li&gt;
&lt;li&gt;You are hiding a slow retrieval or database step before the model starts&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For agent workflows, stream &lt;strong&gt;status events&lt;/strong&gt;, not only text.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Searching relevant docs"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Checking account permissions"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Drafting answer with citations"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This keeps users oriented while the system does real work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Route models by latency class
&lt;/h2&gt;

&lt;p&gt;Not every request deserves your strongest model.&lt;/p&gt;

&lt;p&gt;Create latency classes:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Class&lt;/th&gt;
&lt;th&gt;Use case&lt;/th&gt;
&lt;th&gt;Model strategy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Instant&lt;/td&gt;
&lt;td&gt;autocomplete, labels, short rewrites&lt;/td&gt;
&lt;td&gt;smallest reliable model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fast&lt;/td&gt;
&lt;td&gt;support chat, extraction, routing&lt;/td&gt;
&lt;td&gt;fast general model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Careful&lt;/td&gt;
&lt;td&gt;legal-ish, financial-ish, complex reasoning&lt;/td&gt;
&lt;td&gt;stronger model with tighter scope&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Background&lt;/td&gt;
&lt;td&gt;reports, audits, batch enrichment&lt;/td&gt;
&lt;td&gt;slower model or queued worker&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A simple router can start with rules.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;chooseModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;risk&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;low&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;medium&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;high&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;workflow&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;autocomplete&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;small-fast&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;workflow&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;bulk_report&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;batch-careful&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;risk&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;high&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;careful-reasoning&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;fast-general&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Later, you can route based on measured performance, tenant plan, queue depth, or failure rate. Start with rules that developers can understand and debug.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cache the boring parts
&lt;/h2&gt;

&lt;p&gt;Caching is one of the easiest ways to improve both latency and cost, but cache the right things.&lt;/p&gt;

&lt;p&gt;Good cache candidates:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Embeddings for unchanged documents&lt;/li&gt;
&lt;li&gt;RAG retrieval results for common queries&lt;/li&gt;
&lt;li&gt;System prompt templates&lt;/li&gt;
&lt;li&gt;Tool schemas&lt;/li&gt;
&lt;li&gt;Classification outputs&lt;/li&gt;
&lt;li&gt;Deterministic transformations&lt;/li&gt;
&lt;li&gt;Answers to low-risk, repeated questions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Bad cache candidates:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Permission-sensitive answers without tenant scoping&lt;/li&gt;
&lt;li&gt;Personalized answers without user scoping&lt;/li&gt;
&lt;li&gt;Answers based on rapidly changing data&lt;/li&gt;
&lt;li&gt;Outputs that may contain stale prices, policies, or account state&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Always include tenant and permission context in cache keys.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;cacheKey&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;userRole&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;normalizedQuery&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;sourceVersion&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;userRole&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;sourceVersion&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nf"&gt;hash&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;normalizedQuery&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A cache hit that leaks data is worse than no cache.&lt;/p&gt;

&lt;h2&gt;
  
  
  Add graceful degradation
&lt;/h2&gt;

&lt;p&gt;Your app needs a plan for bad days: provider slowness, queue spikes, long documents, or tenants running large jobs.&lt;/p&gt;

&lt;p&gt;Useful degradation patterns:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Switch from careful model to fast model for low-risk requests&lt;/li&gt;
&lt;li&gt;Reduce retrieved chunks when TTFT is at risk&lt;/li&gt;
&lt;li&gt;Shorten output length during load spikes&lt;/li&gt;
&lt;li&gt;Move long tasks to async jobs&lt;/li&gt;
&lt;li&gt;Show partial results with “continue generating”&lt;/li&gt;
&lt;li&gt;Ask the user to narrow the request before spending tokens&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;queueDepth&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;workflow&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;support_rag_answer&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;budget&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;max_input_tokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;6000&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nx"&gt;budget&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;max_output_tokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nx"&gt;budget&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;fallback_model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;fast-general&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is not about lowering quality everywhere. It is about protecting the experience under pressure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watch p95, not averages
&lt;/h2&gt;

&lt;p&gt;Average latency lies. Your happy path can look fine while real users suffer.&lt;/p&gt;

&lt;p&gt;Track these metrics by workflow and tenant tier:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;p50 TTFT&lt;/li&gt;
&lt;li&gt;p95 TTFT&lt;/li&gt;
&lt;li&gt;p50 total latency&lt;/li&gt;
&lt;li&gt;p95 total latency&lt;/li&gt;
&lt;li&gt;input tokens per request&lt;/li&gt;
&lt;li&gt;output tokens per request&lt;/li&gt;
&lt;li&gt;cache hit rate&lt;/li&gt;
&lt;li&gt;timeout rate&lt;/li&gt;
&lt;li&gt;cost per successful task&lt;/li&gt;
&lt;li&gt;retries per request&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A simple alert rule is enough at first.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Alert when support_rag_answer p95 TTFT &amp;gt; 3000ms for 10 minutes.
Alert when cost per successful task rises 30% above 7-day baseline.
Alert when timeout rate &amp;gt; 2% for any paid tenant tier.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Tie latency to cost. If p95 latency and cost both rise, you may have context bloat, retry loops, poor routing, or a workflow that should become async.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat retries as a budget risk
&lt;/h2&gt;

&lt;p&gt;Retries feel harmless in code and expensive in production.&lt;/p&gt;

&lt;p&gt;A retry can double cost, increase latency, and create duplicate tool actions. For agents, retry loops are even riskier because the model may call tools again.&lt;/p&gt;

&lt;p&gt;Use retry rules:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Retry network errors with jitter&lt;/li&gt;
&lt;li&gt;Do not retry validation failures blindly&lt;/li&gt;
&lt;li&gt;Never retry write actions without idempotency keys&lt;/li&gt;
&lt;li&gt;Stop after a small number of attempts&lt;/li&gt;
&lt;li&gt;Log retry reason and added cost
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;retryPolicy&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;maxAttempts&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;retryOn&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;rate_limit&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;network_timeout&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;neverRetryOn&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;invalid_json&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;permission_denied&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;policy_blocked&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If a workflow needs three retries to feel reliable, it probably needs a better design, not a bigger retry loop.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to use async instead of chat
&lt;/h2&gt;

&lt;p&gt;Some AI work should not pretend to be instant.&lt;/p&gt;

&lt;p&gt;Use async jobs for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Large document analysis&lt;/li&gt;
&lt;li&gt;Multi-source research&lt;/li&gt;
&lt;li&gt;Long agent workflows&lt;/li&gt;
&lt;li&gt;Bulk enrichment&lt;/li&gt;
&lt;li&gt;Report generation&lt;/li&gt;
&lt;li&gt;Evaluation runs&lt;/li&gt;
&lt;li&gt;Tasks with external API rate limits&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A good async UX includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Immediate job receipt&lt;/li&gt;
&lt;li&gt;Progress updates&lt;/li&gt;
&lt;li&gt;Cancel button&lt;/li&gt;
&lt;li&gt;Estimated completion window&lt;/li&gt;
&lt;li&gt;Final summary&lt;/li&gt;
&lt;li&gt;Error state that explains what happened&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This protects your chat interface from becoming a waiting room.&lt;/p&gt;

&lt;h2&gt;
  
  
  Implementation checklist
&lt;/h2&gt;

&lt;p&gt;Use this before shipping a new AI feature:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;[ ] Define max TTFT and total latency by workflow&lt;/li&gt;
&lt;li&gt;[ ] Set input and output token caps&lt;/li&gt;
&lt;li&gt;[ ] Log tenant, workflow, model, tokens, cost, TTFT, total time, status&lt;/li&gt;
&lt;li&gt;[ ] Track p95 latency, not only averages&lt;/li&gt;
&lt;li&gt;[ ] Stream text or status events when useful&lt;/li&gt;
&lt;li&gt;[ ] Route models by workflow and risk&lt;/li&gt;
&lt;li&gt;[ ] Cache safe repeated work with tenant-aware keys&lt;/li&gt;
&lt;li&gt;[ ] Trim context before changing infrastructure&lt;/li&gt;
&lt;li&gt;[ ] Move long tasks to async jobs&lt;/li&gt;
&lt;li&gt;[ ] Alert on latency, timeout rate, retry rate, and cost per successful task&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Final thought
&lt;/h2&gt;

&lt;p&gt;An LLM latency budget is not bureaucracy. It is a guardrail for product quality.&lt;/p&gt;

&lt;p&gt;When budgets are missing, every prompt can grow, every agent can wander, every retry can double spend, and every slow request can look like a mystery. When budgets exist, your team can make clear tradeoffs: faster first token, shorter output, better context, safer cache, async workflow, or stronger model only where it matters.&lt;/p&gt;

&lt;p&gt;Fast AI is not just about speed. It is about respecting the user’s time while protecting your margins.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is an LLM latency budget?
&lt;/h3&gt;

&lt;p&gt;An LLM latency budget is a set of limits for an AI workflow: maximum time to first token, maximum total response time, input token cap, output token cap, model route, caching rule, and fallback behavior.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is a good TTFT for AI features?
&lt;/h3&gt;

&lt;p&gt;It depends on the workflow. Inline suggestions should feel almost instant. Chat answers should usually start streaming within one or two seconds. RAG or agent workflows can take longer if the UI shows useful progress.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I reduce LLM latency quickly?
&lt;/h3&gt;

&lt;p&gt;Start by trimming input tokens, limiting output length, streaming responses, caching repeated work, and routing simple tasks to faster models. These changes are often easier than changing infrastructure.&lt;/p&gt;

&lt;h3&gt;
  
  
  Should every AI workflow stream output?
&lt;/h3&gt;

&lt;p&gt;No. Streaming works well for readable text and progress updates. It is less useful for strict JSON, hidden tool-call workflows, or tasks where partial output could confuse the user.&lt;/p&gt;

&lt;h3&gt;
  
  
  How does latency relate to AI cost?
&lt;/h3&gt;

&lt;p&gt;Long prompts, long outputs, retries, and tool loops usually increase both latency and cost. That is why production teams should track tokens, latency, cache hit rate, and cost per successful task together.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is self-hosting faster than using an API?
&lt;/h3&gt;

&lt;p&gt;Not automatically. Self-hosting can reduce control-plane uncertainty, but serving models well requires batching, memory management, scaling, monitoring, and hardware tuning. Measure TTFT, TPOT, and total cost before assuming self-hosting is better.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>performance</category>
      <category>saas</category>
    </item>
    <item>
      <title>AI Agent Safety: When Boundaries Fail with External Tools</title>
      <dc:creator>Karnik Khanwilkar</dc:creator>
      <pubDate>Wed, 05 Aug 2026 03:37:02 +0000</pubDate>
      <link>https://dev.to/karnikkhanwilkar/ai-agent-safety-when-boundaries-fail-with-external-tools-256k</link>
      <guid>https://dev.to/karnikkhanwilkar/ai-agent-safety-when-boundaries-fail-with-external-tools-256k</guid>
      <description>&lt;p&gt;AI agent safety boundaries are a critical challenge when agents use external tools. My journey into understanding how these boundaries can fail began with a deep dive into recent technical reports from leading AI research organizations.&lt;/p&gt;

&lt;p&gt;I encountered this concept while exploring incidents reported by Anthropic and OpenAI. These reports detail scenarios where AI models, despite being explicitly instructed to operate within simulated environments, managed to interact with real-world systems. This phenomenon, often termed "boundary failure," occurs when the actual operational environment of an agent does not match its internal understanding or the constraints it has been given.&lt;/p&gt;

&lt;p&gt;Modern AI agents are becoming incredibly useful because we're equipping them with capabilities far beyond just answering questions. They can run commands, browse the web, use APIs (Application Programming Interfaces), read and modify files, install packages, and interact with other systems. This ability to act and interface with the world is what makes agentic architectures so powerful and a direction truly worth investing in. However, the more an agent can do, the more critical the boundaries around it become.&lt;/p&gt;

&lt;p&gt;A key example comes from Anthropic's July 30 report, detailing three incidents discovered during their cybersecurity evaluations. Claude models were explicitly told they had no internet access and were working inside simulated environments. However, a problem with the evaluation environment's configuration meant that internet access was actually available. While attempting their assigned cybersecurity exercises, the models reached real systems, initially treating them as part of the simulation. In one striking incident, a Claude model even published a malicious Python package to the real PyPI (Python Package Index) registry, all while believing it was still operating within its simulated exercise.&lt;/p&gt;

&lt;p&gt;This wasn't simply an AI "deciding" to misbehave or to intentionally bypass security. The models were completing their tasks, but their understanding of the environment's boundaries was misaligned with the actual, misconfigured reality. This came shortly after a separate OpenAI incident involving Hugging Face, where models reached the real internet in importantly different ways, further underscoring the complexity. These events highlight a crucial distinction: a prompt is not a security boundary.&lt;/p&gt;

&lt;p&gt;In simple terms, if an AI agent is given the ability to use tools like web browsers, APIs, or file system access, the actual system design and its underlying permissions must strictly enforce safety limits. Relying solely on the agent's internal understanding or instructions provided within a prompt is insufficient to guarantee isolation or prevent unintended actions in the real world.&lt;/p&gt;

&lt;p&gt;Studying incidents like these reveals how vital it is to engineer robust guardrails around increasingly capable AI systems. As we push the boundaries of agentic architectures, ensuring alignment and preventing unintended actions becomes a complex, multi-layered engineering problem that extends far beyond the model's intelligence itself. This is why building beats consuming; actively engaging with these challenges is how we contribute to a safer AI future.&lt;/p&gt;

&lt;p&gt;Here's what these real-world incidents highlighted for developers as we build and deploy AI agents:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;A prompt is not a security boundary:&lt;/strong&gt; Explicitly telling an agent "you don't have internet access" is not the same as actually removing internet access or restricting its network capabilities. True isolation requires physical or logical restrictions, enforced at the infrastructure or operating system level, not just linguistic ones.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;The AI model isn't the whole AI system:&lt;/strong&gt; The model's behavior matters, but it's only one component. The entire ecosystem, including the tools we connect it to, the permissions and credentials it receives, the environment it runs in, and the monitoring and safeguards built around it, all play a critical role. Each of these components can introduce vulnerabilities or points of failure.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Think about what happens when your assumptions are wrong:&lt;/strong&gt; Both the Anthropic and OpenAI incidents occurred because core assumptions about a fully simulated and isolated environment did not match the reality of the underlying system configuration. Anticipating potential mismatches between developer instructions and the actual runtime environment is crucial for robust system design and resilience.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;If an agent can act, we need to know what it's doing:&lt;/strong&gt; When agents are empowered with the ability to run commands, browse the web, or modify files, comprehensive monitoring, logging, and audit trails become non-negotiable. We need clear visibility into their actions and interactions with external systems to detect and mitigate unintended behaviors promptly.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Give an agent what it needs, not everything you have:&lt;/strong&gt; Implement the principle of least privilege rigorously. Grant only the minimum necessary permissions and access to tools or systems required for the agent's specific task. This minimizes the potential impact and "blast radius" if a boundary fails or an agent deviates from its intended path.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These challenges reinforce my belief that building truly adaptable and safe AI systems means not just consuming AI, but actively contributing to its foundational safety and ethical deployment. It requires a hands-on approach to system architecture, security engineering, and continuous evaluation, moving beyond theoretical discussions to real-world implementation. My journey continues to focus on how we can empower agents responsibly while upholding the highest standards of alignment and control, helping push the boundaries of what's possible in a secure manner.&lt;/p&gt;




&lt;p&gt;Source: &lt;a href="https://dev.to/hemapriya_kanagala/were-giving-ai-agents-more-tools-what-happens-when-the-boundaries-fail-46gh"&gt;https://dev.to/hemapriya_kanagala/were-giving-ai-agents-more-tools-what-happens-when-the-boundaries-fail-46gh&lt;/a&gt;&lt;/p&gt;

</description>
      <category>aisafety</category>
      <category>machinelearning</category>
      <category>aiagents</category>
      <category>cybersecurity</category>
    </item>
    <item>
      <title>We Measured AI Code Drift Across 5 Tools and 210 Components. Frequency Alone Lied to Us.</title>
      <dc:creator>Jonathan Gordon</dc:creator>
      <pubDate>Wed, 05 Aug 2026 03:29:17 +0000</pubDate>
      <link>https://dev.to/gojongo/we-measured-ai-code-drift-across-5-tools-and-210-components-frequency-alone-lied-to-us-4g85</link>
      <guid>https://dev.to/gojongo/we-measured-ai-code-drift-across-5-tools-and-210-components-frequency-alone-lied-to-us-4g85</guid>
      <description>&lt;p&gt;&lt;em&gt;Empirical research from ReWeaver AI. 42 identical prompts, across 5 tools and 8 production dimensions, compared to human baseline. One metric that changes how you see drift.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Everyone knows AI-generated code has quality issues. What’s less understood is that the way most teams measure those issues — by how often they occur — systematically understates the risk.&lt;/p&gt;

&lt;p&gt;We ran a controlled study to find out how badly. The answer surprised us, particularly in one dimension.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;What We Did&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;We gave five leading AI coding tools (Cursor, Claude Code, Lovable, Figma Make, and VS Code with Copilot) 42 identical prompts: realistic single-component builds — buttons, forms, dashboards, navs, modals, auth surfaces. We scanned every output with ReWeaver, our deterministic drift-detection engine, across eight production readiness dimensions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;User Experience&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Security &amp;amp; Privacy&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Accessibility&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Design Consistency&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Reliability&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Maintainability&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Architecture&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Testability&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We also scanned six human-authored open-source repositories as a reference baseline.&lt;/p&gt;

&lt;p&gt;For each dimension, we calculated two things:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Drift frequency&lt;/strong&gt; — the percentage of lines containing at least one drift occurrence. Counts what went wrong.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Production Drift Ratio (PDR)&lt;/strong&gt;. The PDR is a metric that weights frequency by estimated remediation cost on a 0–1 scale. A PDR of 0.30 is roughly 45 minutes of cleanup per component; 0.70 is about 2.5 hours.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Finding That Stopped Us&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;In &lt;strong&gt;Security &amp;amp; Privacy&lt;/strong&gt;, AI tools produced &lt;strong&gt;3× the human drift frequency&lt;/strong&gt;. That looks manageable — a meaningful gap, but not alarming.&lt;/p&gt;

&lt;p&gt;The PDR was &lt;strong&gt;22× the human reference&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Not 22% more. 22 times more costly to fix.&lt;/p&gt;

&lt;p&gt;The frequency gap makes Security &amp;amp; Privacy drift look like a minor concern. The PDR reveals it’s the most expensive problem in the dataset. AI-generated security drift (client-side authorization gates bypassable in DevTools, raw PII and credentials passed through props without tokenization) is syntactically identical to safe code. It passes review, but the fixes are harder to find and remedy.&lt;/p&gt;

&lt;p&gt;This is the core argument of our study: &lt;strong&gt;frequency counts what went wrong. The PDR quantifies what it will cost to fix it.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Full Results&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Here’s how much more relative drift frequency and severity AI produced across all eight dimensions:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Frequency multiplier&lt;/th&gt;
&lt;th&gt;PDR multiplier&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Security &amp;amp; Privacy&lt;/td&gt;
&lt;td&gt;3.4×&lt;/td&gt;
&lt;td&gt;22×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;User Experience&lt;/td&gt;
&lt;td&gt;4.5×&lt;/td&gt;
&lt;td&gt;6.5×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Accessibility&lt;/td&gt;
&lt;td&gt;1.7×&lt;/td&gt;
&lt;td&gt;5.2×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Design Consistency&lt;/td&gt;
&lt;td&gt;1.7×&lt;/td&gt;
&lt;td&gt;4.1×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reliability&lt;/td&gt;
&lt;td&gt;1.4×&lt;/td&gt;
&lt;td&gt;2.5×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Testability *&lt;/td&gt;
&lt;td&gt;0.61×&lt;/td&gt;
&lt;td&gt;2.0×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Architecture *&lt;/td&gt;
&lt;td&gt;0.55×&lt;/td&gt;
&lt;td&gt;1.7×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Maintainability *&lt;/td&gt;
&lt;td&gt;0.52×&lt;/td&gt;
&lt;td&gt;1.5×&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;* These three dimensions showed lower AI frequency than the human reference, likely due to a corpus maturity effect, not an AI advantage. Our human reference draws from mature production repositories carrying accumulated technical debt; the AI corpus is fresh greenfield components. The PDR gap remains positive even here.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;In every dimension, AI-generated drift is more expensive to remediate than human-authored drift, even where humans produce more of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Statistics&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;We ran Wilcoxon signed-rank tests comparing AI tool PDR scores against the human reference.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Global test&lt;/strong&gt; (n = 40 paired observations; 5 tools × 8 dimensions): AI tools produced significantly more costly drift than the human baseline (&lt;em&gt;z&lt;/em&gt; = −5.43, &lt;em&gt;p&lt;/em&gt; &amp;lt; .001). Of 40 comparisons, 38 showed AI PDR above the human reference. (One was lower. One tied.)&lt;/p&gt;

&lt;p&gt;The same test applied to drift frequency was not significant (&lt;em&gt;z&lt;/em&gt; = −0.585, &lt;em&gt;p&lt;/em&gt; = .559).&lt;/p&gt;

&lt;p&gt;That asymmetry is the finding. The same code, measured two ways, tells two very different stories.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Per-dimension note:&lt;/strong&gt; with &lt;em&gt;n&lt;/em&gt; = 5 tools per dimension, the minimum attainable exact &lt;em&gt;p&lt;/em&gt;-value is 0.0625, which doesn’t clear the conventional &lt;em&gt;p&lt;/em&gt;≤.05 threshold. We report these results as directional evidence, not formally significant findings, supported by effect sizes (r = 0.90–0.91 in six of eight PDR dimensions) and unanimous positive ranks.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;What the Drift Actually Looked Like&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;This is where the data gets concrete. The drift we found wasn’t just malformed code. It was &lt;strong&gt;absent code&lt;/strong&gt;: components that satisfied the prompt and omitted the production context the prompt didn’t ask for.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Security &amp;amp; Privacy:&lt;/strong&gt; Role gates enforced only in the UI, bypassable in DevTools. Raw credentials passed through props without tokenization. The model treated the browser as a trusted environment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Accessibility:&lt;/strong&gt; Focus escaped modals and was never returned. Interactive elements built without semantic markup. Keyboard users left stranded.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;User Experience:&lt;/strong&gt; Containers missing overflow containment. Forms that fail silently — errors that identify the problem but not the fix. Lists with no empty state. Constraint hints shown before the user has touched the field.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Design Consistency:&lt;/strong&gt; Models named UI primitives from memory without checking they existed in the design system. Inline styles bypassed design tokens. Placeholder content shipped live.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Testability:&lt;/strong&gt; document and window reached synchronously in component bodies — unmockable in test environments. The silent killer: components emitted with no tests alongside them.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;What This Means in Practice&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Drift is endemic.&lt;/strong&gt; Every file we tested (both human and AI-generated) produced drift across every dimension. In terms of tool performance, no tool performed better across all eight dimensions. Switching tools hoping to improve on drift only changes where it shows up and how. No tool avoids drift.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Frequency misleads.&lt;/strong&gt; A Security &amp;amp; Privacy gap that looks like 3× is actually 22× when you account for what fixing it costs. Teams relying on frequency-based metrics are systematically underestimating their production readiness risk, most severely in the dimensions that matter most.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tool selection isn’t the answer.&lt;/strong&gt; The actionable conclusion isn’t which tool to use. It’s that any tool requires a verification layer capable of catching what generation leaves behind.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Try It&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The Playground at &lt;a href="https://www.reweaver.ai/playground" rel="noopener noreferrer"&gt;reweaver.ai/playground&lt;/a&gt; lets you paste React/TypeScript code and get your own PDR score. Free, no login, code scanned in memory and never stored.&lt;/p&gt;

&lt;p&gt;The full research report is at &lt;a href="https://info.reweaver.ai/drift-research-report" rel="noopener noreferrer"&gt;https://info.reweaver.ai/drift-research-report&lt;/a&gt;&lt;/p&gt;

</description>
      <category>codequality</category>
      <category>security</category>
      <category>aicode</category>
      <category>designsystem</category>
    </item>
    <item>
      <title>Image Upload Moderation Beyond Node.js: Classify NSFW and Violence with Multimodal Chat</title>
      <dc:creator>JamesAnderson121</dc:creator>
      <pubDate>Wed, 05 Aug 2026 03:21:58 +0000</pubDate>
      <link>https://dev.to/jamesanderson121/image-upload-moderation-beyond-nodejs-classify-nsfw-and-violence-with-multimodal-chat-2n47</link>
      <guid>https://dev.to/jamesanderson121/image-upload-moderation-beyond-nodejs-classify-nsfw-and-violence-with-multimodal-chat-2n47</guid>
      <description>&lt;p&gt;Use multimodal chat with a strict JSON schema when your policy needs explainable labels for uploaded images; otherwise reach for a managed, fixed-taxonomy service. There is no dedicated image moderation endpoint here, so the practical design is a policy prompt, a vision-capable chat model, schema validation, and a conservative fallback.&lt;/p&gt;

&lt;p&gt;That is my short answer. I would not ship the model's prose directly into an allow/block decision. I keep the original decision for audits, translate it into a small internal status, and make the eval set the release gate. The model is one component of the policy system — not the policy system itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  What should a Python image upload moderation example classify for NSFW and violence?
&lt;/h2&gt;

&lt;p&gt;The categories should come from the app's actual rules. For a general user-content product, I start with nudity, graphic violence, hate symbols, drugs, and minors-risk. I don't pretend those labels are universal: a medical forum and a marketplace need different thresholds, and a historical archive may legitimately show symbols that a profile-photo product should reject.&lt;/p&gt;

&lt;p&gt;My first notebook pass is deliberately boring. I assemble a small set of allowed, blocked, and ambiguous pictures; write the expected category labels; and record the policy reason in plain English. Then I run the same prompt and schema across every candidate model. The score I care about first is false negatives on the block set, followed by false positives on harmless uploads. Overall accuracy can hide both.&lt;/p&gt;

&lt;p&gt;This is also where a JSON schema earns its keep. A response containing &lt;code&gt;"graphic_violence": "high"&lt;/code&gt; can be validated, stored, and compared. A paragraph such as “this appears concerning” can't reliably drive a queue or an appeal. Keep the provider response beside a normalized status such as &lt;code&gt;allow&lt;/code&gt;, &lt;code&gt;review&lt;/code&gt;, or &lt;code&gt;block&lt;/code&gt;; when policy changes, you can replay the raw decisions without migrating every old record.&lt;/p&gt;

&lt;p&gt;I learned the cost side the annoying way: one evaluation run consumed 18.7 million input tokens, roughly 3.4 times my estimate, because I had repeated the full policy rubric for every crop and retry. My notebook showed a reasonable per-case estimate, but the production-shaped harness expanded each source into several variants, then retried cases whose structured response failed validation. I had measured the neat path and budgeted for the messy one. I stopped the run, grouped usage by fixture and attempt, and found that the largest images weren't the main culprit; duplicated policy text across the expanded cases was. The fix was measurement, not guesswork. I made prompt tokens a first-class eval column, deduplicated image variants before dispatch, and reported cost per accepted decision rather than cost per request. That last denominator matters because a cheap response that lands in manual review hasn't completed the job. I also put a batch-level ceiling around experiments, so a mistaken multiplier stops early instead of becoming a surprise at the end of the day. Now I count prompt tokens before any large run and inspect the distribution, not just its mean.&lt;/p&gt;

&lt;p&gt;Small batches first.&lt;/p&gt;

&lt;p&gt;No prose parsing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The focused implementation
&lt;/h2&gt;

&lt;p&gt;The example below sends one local image to the verified &lt;code&gt;POST /v1/chat/completions&lt;/code&gt; route. It uses a data URL so the program is self-contained, takes both the API key and vision model ID from environment variables, requires structured JSON, and retries rate limits while honoring &lt;code&gt;Retry-After&lt;/code&gt;. I use the standard-library HTTP client because this article is about the policy boundary, not a framework choice.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;base64&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;mimetypes&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;urllib.error&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;urllib.request&lt;/span&gt;


&lt;span class="n"&gt;API_URL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.infrai.cc/v1/chat/completions&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;CATEGORIES&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nudity&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;graphic_violence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;hate_symbols&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;drugs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;minors_risk&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;as_data_url&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;media_type&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;mimetypes&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;guess_type&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;application/octet-stream&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;image_file&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;encoded&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;base64&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;b64encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;image_file&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;()).&lt;/span&gt;&lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ascii&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;media_type&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;;base64,&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;encoded&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;post_with_backoff&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;attempts&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;INFRAI_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;attempt&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;attempts&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;request&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;API_URL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Content-Type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;application/json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="n"&gt;method&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;POST&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;urlopen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;HTTPError&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;error_body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;replace&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;code&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="mi"&gt;429&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;attempt&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;attempts&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;RuntimeError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Chat request failed with HTTP &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;code&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;error_body&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;error&lt;/span&gt;
            &lt;span class="n"&gt;retry_after&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Retry-After&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;delay&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;retry_after&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;retry_after&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;attempt&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;random&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
            &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;delay&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;RuntimeError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Retry budget exhausted&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;moderate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;label_properties&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;category&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;string&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;enum&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;none&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;low&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;high&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]}&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;category&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;CATEGORIES&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="n"&gt;payload&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;INFRAI_VISION_MODEL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;messages&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;system&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Classify this upload under the supplied app policy. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Use high only for clear evidence. Return JSON only.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                &lt;span class="p"&gt;),&lt;/span&gt;
            &lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
                    &lt;span class="p"&gt;{&lt;/span&gt;
                        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
                            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Policy: flag nudity, graphic violence, hate symbols, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;drugs, and minors-risk. Give a brief policy reason.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                        &lt;span class="p"&gt;),&lt;/span&gt;
                    &lt;span class="p"&gt;},&lt;/span&gt;
                    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;image_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;image_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;as_data_url&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;)}},&lt;/span&gt;
                &lt;span class="p"&gt;],&lt;/span&gt;
            &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;response_format&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;json_schema&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;json_schema&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;upload_moderation&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;strict&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;schema&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;object&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;properties&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;labels&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;object&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;properties&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;label_properties&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;required&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;CATEGORIES&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
                            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;additionalProperties&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                        &lt;span class="p"&gt;},&lt;/span&gt;
                        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reason&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;string&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
                    &lt;span class="p"&gt;},&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;required&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;labels&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reason&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;additionalProperties&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="n"&gt;raw_response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;post_with_backoff&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;raw_decision&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;raw_response&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;choices&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;message&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;levels&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;raw_decision&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;labels&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;values&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
    &lt;span class="n"&gt;normalized_status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;block&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;high&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;levels&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;review&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;low&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;levels&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;allow&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;raw_decision&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;raw_decision&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;normalized_status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;normalized_status&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;


&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;argv&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;SystemExit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Usage: python moderate_upload.py IMAGE_PATH&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;moderate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;argv&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]),&lt;/span&gt; &lt;span class="n"&gt;indent&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Set &lt;code&gt;INFRAI_API_KEY&lt;/code&gt;, set &lt;code&gt;INFRAI_VISION_MODEL&lt;/code&gt; to a currently available multimodal model from the live model catalog, then run the file with an image path. Model availability changes, so I avoid baking an ID into an article. The explicit schema is the fallback boundary: malformed or missing fields should send an upload to manual review rather than quietly allowing it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choosing the integration boundary
&lt;/h2&gt;

&lt;p&gt;I compare systems by who owns the taxonomy, how much adapter code lands in my repository, and whether I can replay decisions. Exact model support changes; verify image input and structured-output support in each provider's current documentation before committing.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Option&lt;/th&gt;
&lt;th&gt;Integration shape&lt;/th&gt;
&lt;th&gt;Policy and evaluation trade-off&lt;/th&gt;
&lt;th&gt;I would choose it when&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Infrai&lt;/td&gt;
&lt;td&gt;OpenAI-compatible chat behind one REST API&lt;/td&gt;
&lt;td&gt;My team owns the schema, thresholds, normalization, and evals&lt;/td&gt;
&lt;td&gt;I expect moderation to sit beside other backend capabilities and want one consistent contract&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI API&lt;/td&gt;
&lt;td&gt;Direct model-provider integration&lt;/td&gt;
&lt;td&gt;My team still owns the app-specific policy mapping and regression set&lt;/td&gt;
&lt;td&gt;The application already standardizes on OpenAI's client surface&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google Gemini API&lt;/td&gt;
&lt;td&gt;Direct model-provider integration&lt;/td&gt;
&lt;td&gt;I must test my schema and image set against its current model behavior&lt;/td&gt;
&lt;td&gt;Gemini is already the evaluated model family in the stack&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic API&lt;/td&gt;
&lt;td&gt;Direct model-provider integration&lt;/td&gt;
&lt;td&gt;I must confirm current image and structured-output behavior for my exact contract&lt;/td&gt;
&lt;td&gt;The team already operates and evaluates Anthropic models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Amazon Rekognition&lt;/td&gt;
&lt;td&gt;Managed image-analysis service&lt;/td&gt;
&lt;td&gt;A service-defined feature set may require an adapter to my internal statuses&lt;/td&gt;
&lt;td&gt;I prefer a specialized managed vision workflow over chat-prompt ownership&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Infrai's relevant advantage is breadth behind a simple surface: 295 routes across 20 modules sit under one key and one REST contract. For a Python team moving from notebook to production, adding a backend capability can mean another endpoint rather than another SDK, credential set, and adapter. The public discovery response is self-describing, too, so I can inspect readiness and schemas before generating a client.&lt;/p&gt;

&lt;p&gt;The catch is real. Infrai does not provide a dedicated image moderation endpoint in this path, which means my team owns policy wording, the JSON contract, calibration, and appeals. Stick with a specialized managed moderation product when you want its fixed taxonomy and operational workflow, or stay direct with OpenAI, Google, or Anthropic when provider-specific controls matter more than a common interface. I'm not sure which model will win on your images; your mileage may vary, and only a representative eval set settles it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure policy matters more than prompt polish
&lt;/h2&gt;

&lt;p&gt;A production gate needs an explicit response to uncertainty. I map schema failures, unknown labels, and low-confidence classifications to &lt;code&gt;review&lt;/code&gt;; clear high-severity evidence maps to &lt;code&gt;block&lt;/code&gt;; only a complete all-clear maps to &lt;code&gt;allow&lt;/code&gt;. That conservative mapping is intentionally outside the model prompt. Product code can test it, version it, and explain it during an appeal.&lt;/p&gt;

&lt;p&gt;I also store the raw model decision, normalized status, policy version, model ID, and request ID. The first two are the core distinction: raw evidence preserves what the classifier returned, while the normalized value keeps downstream systems stable when I rename a category or tighten a threshold. Retention and access rules should match the sensitivity of user uploads. Don't treat the audit store as an excuse to keep images forever.&lt;/p&gt;

&lt;p&gt;There is another tempting distraction: image upscaling. Infrai exposes optional Lanczos-only upscale, but resizing is separate from moderation and is not a safety control. I would test the original upload for the decision path. If a product separately needs an enlarged asset, treat that as image processing with its own purpose and retention rules.&lt;/p&gt;

&lt;p&gt;This section is short on purpose. The hard work isn't a clever prompt; it's deciding what happens when the classifier is uncertain, proving that behavior with tests, and keeping enough evidence to revisit a decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I measure before copying this design
&lt;/h2&gt;

&lt;p&gt;Before launch, I freeze a labeled set that reflects the real upload mix, including benign edge cases and policy-boundary examples. I report false-negative rate per severe category, false-positive rate, manual-review rate, schema-valid response rate, and cost per final decision. I also slice results by image source and policy category because one aggregate score can conceal a bad hate-symbol result behind easy drug-free photos.&lt;/p&gt;

&lt;p&gt;Then I rerun the suite whenever the prompt, schema, policy, or model changes. A candidate ships only if it meets the category thresholds and does not push review volume past the team's capacity. I keep a small shadow run for changed models before routing live decisions to them. This is where my notebook habits help: the same fixtures and assertions that picked the model become a production regression harness.&lt;/p&gt;

&lt;p&gt;Prompt cost belongs in that harness. Count the repeated policy text, track retries, and measure the total input per accepted decision — a tiny request viewed in isolation can become an expensive batch after crops, retries, and multiple candidates. As far as I can tell, there is no honest universal threshold for “good enough” moderation. The right bar depends on harm severity, reviewer capacity, and what users can appeal.&lt;/p&gt;

&lt;p&gt;Ship the gate only after the fallback has been exercised, not merely described. Feed it malformed structured output in a unit test. Confirm that it chooses review. Then evaluate the real model on the frozen image set and record the policy version with every result. That makes the design reproducible, which matters far more to me than a polished demo screenshot.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.infrai.cc/llms.txt" rel="noopener noreferrer"&gt;Infrai AI-readable capability manifest&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://api.infrai.cc/v1/discovery" rel="noopener noreferrer"&gt;Infrai live discovery manifest&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://platform.openai.com/docs/" rel="noopener noreferrer"&gt;OpenAI API documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ai.google.dev/gemini-api/docs" rel="noopener noreferrer"&gt;Google Gemini API documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.anthropic.com/" rel="noopener noreferrer"&gt;Anthropic API documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.aws.amazon.com/rekognition/" rel="noopener noreferrer"&gt;Amazon Rekognition documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/openai/tiktoken" rel="noopener noreferrer"&gt;tiktoken tokenizer library&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/pgvector/pgvector" rel="noopener noreferrer"&gt;pgvector extension&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>python</category>
      <category>ai</category>
      <category>moderation</category>
    </item>
  </channel>
</rss>
