Unveiledby Anthropic on September 1, 2026, Claude Fable 5.1 and Claude Mythos 5.1 are based on the same model but do not provide access to the same sensitive use cases. Fable is widely available with enhanced safeguards, while Mythos is reserved for verified professionals in the fields of cybersecurity and the life sciences. The most significant innovation, therefore, does not lie in two levels of intelligence, but in a distribution of capabilities based on identity, usage context, and risk.
Two models based on the same design
Gemini 3.8 Flash succeeds Gemini 3.7 Flash three weeks after its launch. Google intends it for long-running software engineering, autonomous agents, and complex business workflows. The documentation states that it is based on Gemini 3.7 Flash and retains a context window of up to 1,048,576 tokens, with a maximum output of 65,536 tokens. Inputs can combine text, images, video, audio, and PDFs, while the documented output remains text-based.[1][2][3]
Gemini 3.8 Flash Cyber uses the same core intelligence but is tailored with specialized features and access rules designed for cyber defense. Google describes it as capable of exploring complex repositories, searching for vulnerabilities, and generating patches. It is not available through consumer channels. Access is provided through the Fairwind program, which is aimed primarily at public authorities, operators of critical infrastructure, and technology platforms deemed essential.[1][4]
This separation should be understood as an architecture for allocating capabilities. The general-purpose model retains safeguards against malicious uses in cybersecurity. The Cyber variant applies more permissive protections for sensitive defensive tasks, in exchange for organization screening, enhanced authentication, and access monitoring. The innovation therefore relates as much to the model’s governance as to its technical performance.[1][4]
An evolution driven primarily by agents
Gemini 3.8 Flash is not a completely separate new family of models. The official documentation states that it is based on Gemini 3.7 Flash. Google attributes the announced gains to training improvements—particularly in cybersecurity—and to long agentic loops used to evaluate and then refine the system. This continuity suggests a significant evolution in post-training and orchestration, rather than a demonstrated architectural break.[1][2]
The most noticeable change is persistence. When working on a difficult task, the model may perform more intermediate steps, call tools multiple times, and check its work before responding. Developers can choose a low, medium, or high level of reasoning. The medium level is applied by default. The high setting can improve certain complex tasks, but it also increases the number of tokens, latency, and the risk that an early error will propagate throughout the rest of the trajectory.[3]
The unit price, therefore, does not fully reflect an agent’s cost. Through December 31, 2026, Google has announced a rate of $0.75 per million input tokens and $3.75 per million output tokens, then $1.50 and $7.50, respectively, starting January 1, 2027. These prices are identical to those of Gemini 3.7 Flash during the launch period, but a job that involves more rounds, outputs, and tool calls may cost more overall.[1][3]
Executive MBA in AI & Business Transformation
The MBA Redesigned for the Age of AI. For experienced executives who want to lead the transformation of their organizations. Paris, Nice, and Dubai.
The system relies on long loops and tools
A conversational model typically generates a response based on a request. An agent receives a goal, selects an action, observes the result, and adjusts its next steps accordingly. In a software repository, this loop may include searching for files, analyzing dependencies, modifying code, running tests, and refining an initial solution. Gemini 3.8 Flash supports function calls, code execution, file searches, structured output, and URL context. The computer usage feature is available as a preview.[3]
Flash Cyber applies this logic to a security chain. The agent can receive a vulnerable code sample, explore execution paths, attempt to reproduce the behavior, pinpoint the cause, and propose a patch. Detecting a weakness and fixing it are two distinct tasks. A patch may eliminate the tested symptom while leaving the root cause intact, break a function, or create a new vulnerability. Security testing, functional testing, and human review therefore remain necessary.
Google also notes that Flash Cyber can be used with CodeMender, its agent specialized in fixing vulnerabilities. The model then provides the reasoning and generation capabilities, while the agent-based environment organizes the tools, tests, and validation. This distinction is important: a score rarely measures the model in isolation. It also depends on the test suite, permissions, the data provided, and the number of allowed attempts.[4]
Technology Framework
| Capacity or Constraint | What You Need to Know |
|---|---|
| Base | Gemini 3.8 Flash is based on Gemini 3.7 Flash. Google attributes the performance gains to training, agentic loops, and cyber specialization. |
| Inputs and Outputs | Input formats: text, images, video, audio, and PDF. Text output as documented in the API documentation. |
| Context | Up to 1,048,576 tokens can be entered. A large window does not guarantee that each piece of information will be used uniformly. |
| Maximum output | 65,536 tokens per response. An agent-based mission may accumulate multiple responses and exceed this volume over the course of the entire trajectory. |
| Reasoning | Three levels are available: low, medium, and high. The medium level is set by default, and the low level is not supported. |
| Tools | Function calls, code execution, file searches, web searches, structured output, URL context, and computer usage in the preview version. |
| API Pricing | Through December 31, 2026, $0.75 per million incoming tokens and $3.75 per million outgoing tokens. Standard rates will double effective January 1, 2027. |
| Cyber Access | Flash Cyber is available only to Fairwind's approved partners. Standard access to Gemini does not include access to this version. |
| Known Limitations | Glitches, occasional slowdowns or timeouts, increased resource consumption during periods of high activity, and dependence on the quality of tools and permissions. |
Code and security are becoming chains of action
For development teams, the value of Gemini 3.8 Flash lies in its ability to run longer tasks. The model can assist with a migration, investigate a regression, or modify multiple files based on test results. This continuity reduces the need for context switching, but it requires explicit success criteria. A task completed by the agent is only useful if the tests cover the expected behavior and if the modification adheres to the project’s architecture.
For security teams, Flash Cyber promises to reduce the time between identifying a vulnerability and proposing a fix. Google says it is already using it on its own code. The company reports that the Chrome team achieved 2.6 times more valid fixes than with the best, larger commercial models in its comparison. It also notes that its Cloud Vulnerability Research team found a vulnerability classified as critical in less than two hours, during a search that would typically have taken several months. These results are interesting, but they come from Google and do not describe either the full sample or the false-positive rate.[1]
The organizational benefit will depend primarily on the ability to integrate the agent into an existing security process. A cautious deployment grants it access to the necessary code, executes it in an isolated environment, logs tool calls, and requires approval before any merge or release to production. Direct access to secrets, customer data, or production systems increases risk without being essential for most audits.
What the published results allow us to conclude
The published evaluations support a limited conclusion. Gemini 3.8 Flash outperforms Gemini 3.7 Flash on several software engineering and specialized tasks. Flash Cyber outperforms previous versions of Google in reproducing or discovering vulnerabilities. They do not show that a single model dominates all use cases, nor that an agent can reliably patch a production system without validation.
| Evaluation | Published Comparison | Gemini 3.8 | Read Carefully |
|---|---|---|---|
| DeepSWE v1.1 | 3.7 Flash 65.3% | 73,7 % | The public ranking is rounded to 74% with a confidence interval. The best models are close to each other, and their intervals overlap. |
| Vals Finance Agent v2 | 3.7 Flash 59.0% | 61,4 % | Progress on financial analyst tasks within the approved protocol, without general validation for actual finance. |
| Harvey, Legal Agent | 3.7 Flash 8.8% | 10,0 % | Highest score on the Google chart, but a low overall success rate for complete legal workflows. |
| CyberGym Pass@1 | 3.5 Flash Cyber 77.5% | 86,2 % | High score on the reproduction of known vulnerabilities. The benchmark alone does not measure the full scope of a discovery or the quality of a patch. |
| Internal analysis of 20 languages | 3.7 Flash 58.9% | 71,0 % | Significant improvement in an internal Google test, although independent replication remains limited. |
| CWE-Bench Pass@1 | Fable 5 47.8% | 47,2 % | Similar performance at a lower advertised cost. The score remains below 50% and is subject to benchmark validation. |
DeepSWE provides a useful public verification. As of September 22, 2026, its ranking places Gemini 3.8 Flash at approximately 74%, with an average cost of $2.36, a cumulative output of 143,000 tokens, and 166 steps per task. Gemini 3.7 Flash achieves approximately 65%, at $2.03, 94,000 tokens, and 117 steps. Progress therefore comes with a longer execution time and a higher average cost on this test, despite an identical cost per token.[5]
CyberGym warrants a careful read. The benchmark brings together 1,507 historical vulnerabilities from 188 projects. At its core, the agent receives a description and an unpatched repository, and must then produce a proof-of-concept that replicates the vulnerability. This is a realistic replication task, but it differs from a completely open-ended discovery. The authors also note that the results are submitted by the teams, that the executions are stochastic, and that small deviations may not represent a significant difference.[6]
Patching results require extra caution. Recent work on PatchBench shows that evaluations based on a proof-of-concept test may overestimate the quality of patches. Some agents generate a patch that closely resembles a previously recorded fix or eliminate the symptom without addressing the root cause. In this study, validation limited to the initial test inflates the success rate by a factor of 1.83 on average. The CWE-Bench result should therefore be interpreted as an experimental indicator, not as a guarantee of a successful fix in production.[7]
Controlled access is becoming a feature of the product
Cyber capabilities are dual-use. A system capable of explaining why a vulnerability works possesses some of the knowledge needed to reproduce it. Google addresses this tension with Fairwind. Partners can conduct authorized threat simulations, reverse engineering, and malware analysis for defensive or academic purposes. They must implement single sign-on, phishing-resistant multi-factor authentication, access controls, and usage tracking.[4]
The program prohibits the sharing, resale, or redistribution of access. Google also verifies applicant organizations and restricts internal use to cybersecurity, incident response, or penetration testing teams. These rules reduce certain opportunities for abuse, but they do not guarantee that an approved account will always remain legitimate. An organization may be compromised, an authorized user may exceed their scope of authority, and a defensive request may yield information that can be reused elsewhere.
Governance must therefore continue after admission. It is necessary to define the duration of rights, accessible repositories, usable tools, approval thresholds, and the revocation procedure. Logs must make it possible to reconstruct the data accessed, the commands executed, the proposed changes, and the human approvals. Differentiated access is a security measure, not proof that all data exports are secure.
Organizations must secure the agent before the code
An agent connected to tools encounters unreliable data. A malicious instruction may be hidden in a web page, file, ticket, or code comment in order to divert its path. Google reports a 6% attack success rate for Flash Cyber on the Gray Swan IPI benchmark, where a lower value is preferable. This result indicates greater resilience in the published protocol, though it does not demonstrate immunity. Gray Swan also notes that none of the models tested in its indirect injection competition were fully protected.[8]
The defense strategy is therefore multi-layered. External content must be treated as data; sensitive information must be kept out of context when its presence is not necessary; and tools must adhere to the principle of least privilege. Actions that are irreversible or likely to affect production require approval. Runtime environments must restrict access to the network, files, and identities. Finally, teams must test the entire system, as model security does not automatically cover the application’s connectors, scripts, and policies.
Human supervision is not simply a matter of mechanically clicking to approve something. The reviewer must understand the change, be aware of any missing tests, and be able to reject or correct the patch. Cybersecurity roles are thus evolving toward the evaluation of agent behaviors, the design of access policies, the verification of evidence, and the analysis of incidents generated by the agents themselves.
Costs and impacts must be measured by mission
The price of Gemini 3.8 Flash may make experimentation easier, but an organization must track the cost per completed task. This metric includes tokens, tool calls, computing in isolated environments, failures, retries, and review time. An agent that is cheaper per token can end up being more expensive if it multiplies the number of steps or produces very long outputs. Public data from DeepSWE clearly illustrates this discrepancy between unit price and operational cost.[5]
The same reasoning applies to environmental impact. Google acknowledges that the model may use more tokens on complex tasks, but does not provide data in the document on energy consumption per task or on the carbon footprint specific to Gemini 3.8 Flash. It would therefore be premature to quantify any improvement or deterioration. Organizations can nevertheless reduce unnecessary computations by choosing the appropriate level of effort, limiting loops, reusing the cache, and halting trajectories that are no longer making progress.[3]
The legal framework depends on the use and the sector
In the European Union, the AI Act does not automatically classify a cybersecurity agent as a high-risk system. Classification depends on the actor’s role, the purpose, and the deployment context. A system used as a security component of critical digital infrastructure may be subject to stricter requirements. As of August 2, 2026, the AI Office and national authorities will implement and enforce the applicable provisions of the regulation, with a separate timeline for certain high-risk categories.[9]
The GDPR remains applicable when an employee processes personal data contained in logs, incident tickets, support requests, or customer environments. The organization must limit the data transmitted, establish a legal basis and a retention period, monitor transfers, and document who has access to the logs. Technical secrets, keys, and credentials require specific protection, even when they are not personal data.
The European Cyber Resilience Regulation adds another dimension for products containing digital components. Its reporting requirements regarding actively exploited vulnerabilities and serious incidents have been in effect since September 11, 2026, while the bulk of the regulation will take effect on December 11, 2027. A patch generated by an agent does not replace vulnerability analysis, documentation, or the manufacturer’s obligations.[10]
What to Watch for Now
The first question concerns independent reproduction. The internal findings regarding twenty languages, the patches released for Chrome, and the vulnerability discovered in less than two hours should be documented with sufficient protocols, samples, and false-positive rates. Without these elements, they remain merely findings announced by Google and its partners.
The second question concerns the quality of the fixes. We will need to assess whether the root cause has been addressed, whether functional tests have been successful, whether there are any regressions, how robust the code is against variants of the attack, and how maintainable the code is. A successful proof-of-concept alone does not answer all of these questions.
The third issue is that of access. According to Google, Fairwind has more than 650 partners, but the effectiveness of the governance model will depend on incidents, revocations, audits, and transparency regarding admission criteria. The advantage given to defenders can only be assessed by observing how capabilities spread and how attackers adapt.[4]
Gemini 3.8 Flash and Flash Cyber thus demonstrate the evolution of agent-based AI toward longer and more sensitive tasks. Their true impact will be measured by the reliability of complete trajectories, the cost per mission, the quality of corrections, and organizations’ ability to restrict permissions. In cybersecurity, model performance and the security of its environment have become inseparable.
Learn more
To understand Gemini 3.8 Flash and Flash Cyber within the context of the evolution of code agents, specialized models, and careers in cybersecurity, continue reading these analyses from the aivancity blog.
Sources
[1] Google, September 2, 2026. Introducing Gemini 3.8 Flash and 3.8 Flash Cyber. View source
[2] Google DeepMind, September 2026. Gemini 3.8 Flash Model Card. View source
[3] Google AI for Developers, accessed September 30, 2026. Gemini 3.8 Flash Latest Model Documentation. View source
[4] Google DeepMind, accessed September 30, 2026. Gemini 3.8 Flash Cyber and Fairwind Program. View source
[5] Datacurve, rankings updated on September 22, 2026. DeepSWE v1.1 Leaderboard. View source
[6] Wang, Z. et al., ICLR 2026. CyberGym: Evaluating AI Agents’ Real-World Cybersecurity Capabilities at Scale. View source
[7] Shen, C. et al., September 3, 2026. PatchBench: Evaluating AI Agents for Vulnerability Patching. View source
[8] Gray Swan AI, accessed September 30, 2026. Indirect Prompt Injection Arena and Benchmark Resources. View source
[9] European Commission, accessed September 30, 2026. The Enforcement Framework of the AI Act. View source
[10] EUR-Lex, Regulation (EU) 2024/2847. Cyber Resilience Act and Application Calendar. View source

