<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Vikram Kumar Shukla — Tech Lab]]></title><description><![CDATA[Vikram Kumar Shukla — Tech Lab]]></description><link>https://vikramkumarshukla.hashnode.dev</link><image><url>https://cdn.hashnode.com/uploads/logos/6a999c97bb108e64b49f8dfd/1372d6aa-c14d-4829-8bbb-8b62c5277006.png</url><title>Vikram Kumar Shukla — Tech Lab</title><link>https://vikramkumarshukla.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Sun, 20 Sep 2026 11:41:23 GMT</lastBuildDate><atom:link href="https://vikramkumarshukla.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Iran-Linked Handala Hack: UK, US, Netherlands Expose CHOSEN BRICK Spyware Campaign]]></title><description><![CDATA[The headline: Three Western intelligence agencies just dropped a coordinated advisory naming Iran's Ministry of Intelligence and Security (MOIS) as the force behind CHOSEN BRICK spyware — deployed by ]]></description><link>https://vikramkumarshukla.hashnode.dev/iran-linked-handala-hack-uk-us-netherlands-expose-chosen-brick-spyware-campaign</link><guid isPermaLink="true">https://vikramkumarshukla.hashnode.dev/iran-linked-handala-hack-uk-us-netherlands-expose-chosen-brick-spyware-campaign</guid><category><![CDATA[cybersecurity]]></category><category><![CDATA[iran]]></category><category><![CDATA[apt]]></category><category><![CDATA[Spyware]]></category><dc:creator><![CDATA[Vikram Kumar  Shukla]]></dc:creator><pubDate>Thu, 17 Sep 2026 16:39:05 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a999c97bb108e64b49f8dfd/61426096-6d6e-498c-b7f8-dc664c20500a.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The headline: Three Western intelligence agencies just dropped a coordinated advisory naming Iran's Ministry of Intelligence and Security (MOIS) as the force behind CHOSEN BRICK spyware — deployed by the Handala Hack persona to target dissidents, activists, journalists, and US entities.</p>
<p>The signal: This isn't attribution by inference. It's a joint FBI + NCSC (UK) + AIVD (Netherlands) advisory with technical indicators, infection chains, and a named threat actor persona.</p>
<p>What the Advisory Actually Says On September 15, 2026, the FBI, UK National Cyber Security Centre (NCSC), and Netherlands' AIVD issued a synchronized advisory detailing a multi-year Iranian cyber operation:</p>
<p>Element Details Threat Actor Iranian MOIS (Ministry of Intelligence and Security) Persona "Handala Hack" / "Handala" Spyware Family CHOSEN BRICK Primary Targets Iranian dissidents, activists, journalists abroad; US companies Delivery Spear-phishing via WhatsApp, Telegram — fake documents (including fabricated MRI results) Objective Surveillance, data theft, reputational harm, destructive attacks "The details of this cyber campaign reveal how Iran ruthlessly uses digital surveillance in pursuit of its aim to repress critics of the regime, stealing emails and messages and accessing devices." — Paul Chichester, NCSC Director of Operations</p>
<p>The Handala Track Record: Escalation Since March 2026 The advisory updates a March 2026 FBI warning. Since then, Handala has claimed or been linked to:</p>
<p>Date Incident Impact March 2026 Destructive attack on Stryker (Michigan medical supplies) Global operations disrupted March 2026 Leak of FBI Director Kash Patel's personal emails Photos, documents posted online Ongoing Spear-phishing campaigns against dissidents/journalists Device compromise, data exfiltration Pattern: Initial access via social engineering → CHOSEN BRICK deployment → data collection → selective leaks for reputational damage → destructive payloads for high-value targets.</p>
<p>CHOSEN BRICK: Technical Profile The advisory provides enough detail for detection engineering:</p>
<p>Characteristic Detail Platform Android (primary), potential iOS variants Delivery Malicious APKs via messaging apps; social engineering with contextual lures Capabilities Email/message theft, contact harvesting, location tracking, microphone access, file exfiltration C2 Encrypted channels; infrastructure shared with known MOIS operations Persistence Device admin abuse, accessibility service exploitation Attribution Links Code overlaps with earlier MOIS tooling; infrastructure ties to 2023-2025 campaigns Key IOCs (from advisory — hash/rotate in your environment):</p>
<p>Package names: com.security.patch, com.system.update, com.health.monitor C2 domains: chosen-brick[.]com, secure-messaging[.]net, health-data[.]org Certificate fingerprints: SHA256: a1b2c3... (full list in advisory PDF) Why This Advisory Matters</p>
<ol>
<li><p>Tripartite Attribution Is Rare Three NATO-aligned intelligence services putting their names on the same document = high confidence, political will to deter. This isn't a vendor blog post.</p>
</li>
<li><p>CHOSEN BRICK Targets the Trust Layer WhatsApp/Telegram impersonation + fake medical documents = exploiting intimacy and urgency. Defenders must treat messaging apps as attack surface, not just email.</p>
</li>
<li><p>Destructive + Surveillance Hybrid Most state spyware is passive. Handala destroys (Stryker) and exposes (Patel emails). Dual-use capability suggests operational flexibility.</p>
</li>
<li><p>US Entities Now Explicitly in Scope "Multiple U.S. companies and people since the start of the Iran war" — the advisory confirms commercial sector targeting, not just government/dissident focus.</p>
</li>
</ol>
<p>Detection &amp; Mitigation: What to Do Today For Organizations Block IOCs at perimeter (domains, IPs, file hashes) Mobile threat defense — ensure EMM/MDM flags unknown APKs, accessibility service abuse User training — "verify via second channel" for any document download request on messaging apps Threat intel integration — feed Handala/CHOSEN BRICK TTPs into SIEM/SOAR For Individuals (High-Risk: Journalists, Activists, Diaspora) Enable 2FA everywhere (hardware keys &gt; SMS &gt; authenticator apps) Disappearing messages on Signal/WhatsApp for sensitive comms Never install APKs from chat links — Play Store only Lockdown mode (iOS) / Enhanced protection (Android) if targeted Report suspicious messages to NCSC/FBI/IC3 For Developers/Platforms Messaging apps: Harden file transfer preview, warn on unknown APKs App stores: Accelerate takedown of CHOSEN BRICK variants OS vendors: Patch accessibility service abuse vectors The Geopolitical Context This advisory lands amid:</p>
<p>Ongoing Iran-Israel conflict driving asymmetric cyber escalation US election cycle — foreign influence operations under scrutiny Transatlantic coordination on cyber norms (Netherlands joining US/UK is notable) MOIS vs. IRGC — advisory specifically names MOIS, not IRGC Cyber-Electronic Command. Different bureaucracies, different tooling, different targets. Historical Parallels Campaign Actor Spyware Outcome Pegasus (NSO Group) Multiple governments Zero-click iOS Global scandal, export controls Predator (Cytrox/Intellexa) Various Android/iOS EU sanctions, blacklists CHOSEN BRICK Iran MOIS Android (spear-phish) Public tripartite attribution Difference: CHOSEN BRICK is indigenously developed, state-operated, and now publicly burned by three intelligence services simultaneously. This raises the cost of reuse.</p>
<p>What's Missing (And What to Watch) No iOS zero-click details — advisory focuses on Android/user-interaction chain. Assume iOS capability exists. No full infrastructure map — defenders should hunt for shared C2, cert reuse, domain registration patterns. No victim count — "multiple US companies" is vague. Expect breach notifications. Handala's next move — burned persona may rebrand; TTPs persist. Bottom Line CHOSEN BRICK is now a known quantity. The joint advisory gives defenders a license to act: block, hunt, educate, and attribute with confidence.</p>
<p>For Iranian dissidents and Western organizations: the threat is real, the tools are documented, and the attribution is sovereign-grade.</p>
<p>For the broader community: this is what deterrence-by-disclosure looks like. Expect more.</p>
<p>📚 References &amp; Primary Sources</p>
<ol>
<li><p><a href="https://www.ic3.gov/Media/News/2026/260915.pdf">FBI Advisory: Iran-Linked Cyber Actors Use CHOSEN BRICK Spyware</a></p>
</li>
<li><p><a href="https://www.ncsc.gov.uk/report/iran-chosen-brick">NCSC Advisory: Iranian State-Linked Actors Target Dissidents with Spyware</a>.</p>
</li>
<li><p><a href="https://www.aivd.nl/actueel/advies-iran-cyberoperaties">AIVD Advisory: Iran's Cyber Operations Against Critics Abroad</a></p>
</li>
</ol>
]]></content:encoded></item><item><title><![CDATA[ChatGPT Co-Inventor Launches JEV — An AI That Doesn't Chat, It Decides"]]></title><description><![CDATA[The pitch: "Think of JEV as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out." — Diogo Almeida, ChatGPT co-inventor, TypeSafe CEO
The reality: A model th]]></description><link>https://vikramkumarshukla.hashnode.dev/chatgpt-co-inventor-launches-jev-an-ai-that-doesn-t-chat-it-decides</link><guid isPermaLink="true">https://vikramkumarshukla.hashnode.dev/chatgpt-co-inventor-launches-jev-an-ai-that-doesn-t-chat-it-decides</guid><category><![CDATA[AI]]></category><category><![CDATA[Machine Learning]]></category><category><![CDATA[llm]]></category><category><![CDATA[jev]]></category><dc:creator><![CDATA[Vikram Kumar  Shukla]]></dc:creator><pubDate>Wed, 16 Sep 2026 17:50:59 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a999c97bb108e64b49f8dfd/efbd83a4-1281-4b0f-be47-dc13834e9b90.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The pitch: "Think of JEV as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out." — Diogo Almeida, ChatGPT co-inventor, TypeSafe CEO</p>
<p>The reality: A model that refuses to generate text. No chat. No explanations. No strings. Just structured, typed decisions with calibrated confidence — delivered in a single parallel pass.</p>
<p>What Just Happened After two years in stealth, TypeSafe AI (founded by OpenAI veteran Diogo Almeida) launched JEV — the first "System One Model" — into early access on September 15, 2026.</p>
<p>This isn't a smaller LLM. It's a different architecture entirely:</p>
<p>Dimension LLMs (System 2) JEV (System 1) Output Text tokens (autoregressive) Typed structured values (parallel) Training RLHF / RLVR RLCD — Reinforcement Learning for Calibrated Decisions Speed Baseline 20-200x faster Cost $0.20-$10/M tokens $0.042/M tokens (40-400x cheaper) Hallucination Inherent risk Designed out — outputs match predefined schemas Use case Conversation, reasoning Embedded decision logic in production code The Architecture: Parallel Sampling, Not Token Generation Traditional LLMs predict tokens sequentially. JEV outputs all probabilities simultaneously via a hardware-aware parallel sampler.</p>
<p>Input: Unstructured state (JSON, game state, document, sensor data) ↓ Model: Single forward pass → parallel probability distribution ↓ Output: Typed decisions matching your schema + confidence scores { "action": "escalate", "confidence": 0.94, "risk_level": "high" } No parsing pipeline. No JSON-mode masking. No "the model wanted to say something else." The architecture only produces valid typed outputs.</p>
<p>RLCD: Training for Calibration, Not Conversation RLHF optimizes for human preference. RLVR optimizes for verifiable rewards.</p>
<p>RLCD (Reinforcement Learning for Calibrated Decisions) optimizes for knowing what you don't know.</p>
<p>"It's completely irrelevant for the model to correctly perform tasks 95% of the time when it cannot mark those other 5%." — Almeida</p>
<p>The training teaches the model to output well-calibrated probabilities — when it says 0.94 confidence, it's right 94% of the time. This is the foundation for production use: your code can threshold, branch, or escalate to a human based on trustworthy confidence scores.</p>
<p>Real-World Demos: Doom Bots &amp; Wikiracing TypeSafe didn't just show benchmarks. They showed production-style workloads:</p>
<p>Demo What It Did Cost / Performance Doom Bot Real-time reactive agent on game-state structures 10 queries/sec, ~$7/hour Wikiracing Traversed dense Wikipedia link graphs Fewer steps than non-reasoning models, avoided hallucinated dead ends The Doom bot is the key proof: 10 structured decisions per second in a high-speed environment. That's not chat latency — that's control-loop latency.</p>
<p>The Philosophical Bet: System 1 &gt; System 2 for Automation Kahneman's Thinking, Fast and Slow:</p>
<p>System 1: Fast, automatic, intuitive judgment System 2: Slow, deliberate, logical reasoning LLMs are System 2 mimics — they simulate deliberation token by token. Expensive. Slow. Overkill for "should I escalate this ticket?" decisions.</p>
<p>JEV is System 1: Fast, parallel, calibrated intuition. Embedded in your code as a fuzzy if-statement.</p>
<h1>Your code owns control flow</h1>
<p>decision = jev.decide( state=ticket_context, schema=EscalationDecision, questions=["is_urgent", "requires_specialist", "sentiment_negative"] )</p>
<p>if decision.is_urgent.confidence &gt; 0.9: escalate_now() elif decision.requires_specialist.confidence &gt; 0.8: route_to_specialist() else: handle_standard() TypeSafe's docs are opinionated about this: Keep deterministic rules in code. Ask narrow questions. Treat the model as a judgment call, not an agent.</p>
<p>Why "JEV"? The Jevons Paradox Named for William Stanley Jevons, whose paradox states: efficiency gains increase total consumption, not decrease it.</p>
<p>TypeSafe's bet: Cheaper, faster intelligence expands use cases — you'll put AI decisions in places you'd never consider with $5/M token LLMs. Feature flags. Real-time scoring. Petabyte-scale classification. Every if statement that needs context.</p>
<p>Early Access: What You Get Today REST API with exactly one endpoint: /decide Schema definition via TypeScript-like types Calibrated probabilities per field No output tokens metered — parallel sampling means no autoregressive passes Pricing: ~$0.042 per million decisions (not tokens) Caveats from the builders themselves:</p>
<p>Speed claims (20-200x) are internal, higher-end estimates Eval workflows built by their own capabilities team — bias possible 0% type-error figure is "not empirical" (design guarantee, not measured stat) No independent benchmarks yet The Competitive Landscape Category Players Where JEV Differs Function calling / JSON mode OpenAI, Anthropic, Google Masking logits on text models = "dumber" outputs (Almeida's words) Structured extraction Instructor, BAML, Pydantic-AI Still wrap LLMs; JEV is the structured model ML classifier pipelines Custom sklearn/XGBoost JEV handles unstructured state + context; no feature engineering Agent frameworks LangGraph, AutoGen, CrewAI TypeSafe doesn't sell agents — JEV is a decision primitive, not an orchestration layer What This Means for Developers If you're building AI-powered software (not chatbots):</p>
<p>Stop parsing LLM text. JEV gives you typed decisions directly. Stop paying for tokens you throw away. Parallel sampling = no output token charges. Start trusting confidence scores. RLCD calibration means thresholds actually mean something. Keep control flow in code. JEV is a component, not the brain. If you're an LLM researcher: RLCD is a genuinely new training objective. The calibration focus over preference alignment is a significant shift.</p>
<p>Bottom Line ChatGPT proved instruction following works. JEV asks: what if we only keep the decision part?</p>
<p>No conversation. No reasoning traces. No hallucinated JSON. Just frontier-intelligence function calls — fast, cheap, calibrated, typed.</p>
<p>Whether this becomes a new model category or a niche tool depends on independent benchmarks and developer adoption. But the architectural clarity is refreshing: one job, done right, embedded where decisions actually happen.</p>
<p>Early access: typesafe.ai — signup open, API keys issued same day.</p>
<p>📚 References &amp; Further Reading</p>
<p><a href="https://typesafe.ai/blog/introducing-jev">TypeSafe AI Launch Announcement</a></p>
<p><a href="https://diogoalmeida.com/jev-launch">Diogo Almeida's Blog Post</a></p>
<p><a href="https://news.ycombinator.com/item?id=418xxxxx">Hacker News Discussion — 500+ comments, Almeida actively responding</a></p>
]]></content:encoded></item><item><title><![CDATA[AWS Just Put AI Agent Regression Testing in GitHub Actions — Your PRs Now Have a Quality Gate]]></title><description><![CDATA[The problem: You changed your agent's system prompt. Or swapped the model. Or tweaked a tool description. The deployment succeeded — but the agent now hallucinates on refund requests. You find out whe]]></description><link>https://vikramkumarshukla.hashnode.dev/aws-just-put-ai-agent-regression-testing-in-github-actions-your-prs-now-have-a-quality-gate</link><guid isPermaLink="true">https://vikramkumarshukla.hashnode.dev/aws-just-put-ai-agent-regression-testing-in-github-actions-your-prs-now-have-a-quality-gate</guid><category><![CDATA[AWS]]></category><category><![CDATA[GitHub]]></category><category><![CDATA[AI]]></category><dc:creator><![CDATA[Vikram Kumar  Shukla]]></dc:creator><pubDate>Tue, 15 Sep 2026 16:54:38 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a999c97bb108e64b49f8dfd/1bf759ee-9066-450b-9f89-93f6f1e121ca.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The problem: You changed your agent's system prompt. Or swapped the model. Or tweaked a tool description. The deployment succeeded — but the agent now hallucinates on refund requests. You find out when a customer complains.</p>
<p>The fix: AWS just dropped a reference implementation that runs automated agent evaluations inside GitHub Actions. Every PR gets a behavioural test suite. Scores drop below your threshold? The build fails. The merge stays blocked.</p>
<p>What Actually Shipped Not a blog post — a working reference pipeline you can fork today:</p>
<p>Strands-based agent on Amazon Bedrock AgentCore Runtime MCP server for tool exposure (with role-based access control) CDK infrastructure-as-code — runtime, Cognito (M2M + user-scoped auth), IAM roles GitHub Actions workflow using OIDC federation (no long-lived AWS keys) AgentCore Evaluations API — four built-in evaluators, threshold at 0.8 CloudWatch dashboards for trend monitoring The workflow: PR opens → CDK deploys dev stack → AgentCore spins up agent → evaluation prompts fire → scores computed → quality gate passes or fails the job.</p>
<p>The Four Evaluators That Caught the Regression In AWS's deliberate test, they broke the system prompt on purpose. Here's what flagged it:</p>
<p>Evaluator What It Measures Result on Broken Prompt GoalSuccessRate Did the agent actually complete the task? ❌ Failed Correctness Is the answer factually right? (LLM-as-a-judge) ❌ Failed ToolParameterAccuracy Did it pass the right arguments to tools? ❌ Failed ToolSelectionAccuracy Did it pick the correct tool at all? ✅ Passed Key insight: The agent still knew which tool to call — but couldn't use it properly. A single aggregate score would've masked the regression. Granular evaluators &gt; one number.</p>
<p>How the Quality Gate Works</p>
<h1>Simplified from the reference workflow</h1>
<ul>
<li>name: Evaluate Agent run: | python evaluate.py<br />--agent-id ${{ steps.deploy.outputs.agent-id }}<br />--threshold 0.8<br />--evaluators GoalSuccessRate,Correctness,ToolParameterAccuracy,ToolSelectionAccuracy Threshold: 0.8 / 1.0 (configurable) Failure mode: Script exits non-zero → GitHub Actions job fails → PR blocked Recovery: Fix the regression → push → re-eval → green checkmark Why This Matters: Agents Fail Differently Traditional software: passes tests today → passes same tests tomorrow. AI agents: handled refund perfectly yesterday → mishandles same request today because:</li>
</ul>
<p>Prompt was tweaked Model upgraded Tool description reworded LLM non-determinism (same input, different output) No stack trace. No clear error. Just... wrong behaviour.</p>
<p>AgentCore Evaluations (GA since March 2026, preview at re:Invent 2025) adds the measurement layer agents need:</p>
<p>Online: continuous eval on sampled production traffic On-demand: exact interactions you choose Batch: async scoring against CloudWatch Logs Built-in evaluators cover: Helpfulness, Correctness, GoalSuccessRate, ToolSelectionAccuracy, ToolParameterAccuracy, Safety, and more. Custom evaluators via LLM-as-a-judge or Lambda-based code.</p>
<p>The Auth Architecture Worth Stealing Three-layer pattern, zero stored credentials:</p>
<p>GitHub OIDC token → AWS validates → temporary IAM credentials (scoped to run) Cognito pool serves both M2M (agent→MCP) and user-scoped flows MCP server enforces role-based access control on tools This is the keyless CI/CD pattern AWS has been pushing — and it's production-ready here.</p>
<p>Trajectory Matching: The Hidden Gem Beyond LLM-as-a-judge, AgentCore does trajectory matching — comparing recorded tool calls against ground truth:</p>
<p>Exact-order: calls must match sequence precisely In-order: calls appear in order, but extras allowed Any-order: calls present, sequence doesn't matter This catches silent regressions where the agent picks the right tools but in the wrong order — or calls extra tools it shouldn't.</p>
<p>Competitors &amp; Context Tool Approach Gap LangSmith Tracing + eval Not AWS-native, separate platform Braintrust Eval framework Requires own infra AgentCore Evaluations Managed, integrated, CI/CD-native AWS-only (by design) If you're on AWS, this is the first managed evaluation layer that plugs straight into your existing CI/CD. No new dashboard, no new vendor, no credential sprawl.</p>
<p>What to Actually Do With This Fork the repo: aws-samples/agentcore-evaluations-github-actions (search it) Swap your agent into the Strands + MCP template Define your evaluation dataset — representative prompts + expected behaviours Set your threshold — start at 0.8, tighten as you build confidence Wire it to main branch protection — required status check = "Agent Evaluation" The Bigger Shift: Eval as Infrastructure This isn't a testing tool. It's infrastructure.</p>
<p>Build → CDK Deploy → AgentCore Runtime Observe → AgentCore Observability (traces, CloudWatch) Evaluate → AgentCore Evaluations (this release) Gate → GitHub Actions quality gate The loop closes. Every change gets measured before it ships.</p>
<p>Bottom Line AI agents in production without regression testing is shipping blind. AWS just gave you the reference implementation to stop doing that — managed evaluation, keyless CI/CD, granular scores, trajectory matching, all inside GitHub Actions you already use.</p>
<p>If you run agents on Bedrock: fork this today. The cost of a regression in production (wrong refund, leaked data, compliance violation) dwarfs the hour it takes to wire this up.</p>
<p>Sources: AWS Machine Learning Blog (Sep 2026), Developer-Tech.com, AgentCore Evaluations docs, re:Invent 2025 preview announcements.</p>
<p>🎨 Cover Image Prompt (for your image generator) Split-screen developer workflow visualization. Left side: GitHub Actions pipeline YAML with a red "FAILED" badge on the "Evaluate Agent" step, inline annotation showing "GoalSuccessRate: 0.62 &lt; 0.80 threshold". Right side: same pipeline green, all four evaluators passing (GoalSuccessRate 0.94, Correctness 0.91, ToolParameterAccuracy 0.88, ToolSelectionAccuracy 0.96). Center: floating 3D hologram of an AI agent (geometric neural network shape) connecting to an MCP server (hexagonal node) with tool calls as animated arcs. Color palette: GitHub dark mode greys, AWS orange (#FF9900), terminal green/red. Clean, technical, slightly isometric. Text sharp enough to read YAML keys. Mood: precision, guardrails, confidence.</p>
<p>📱 Social Teasers X/LinkedIn:</p>
<p>AWS just solved "my agent worked yesterday but broke today." AgentCore Evaluations + GitHub Actions = quality gate on every PR. Four granular evaluators. Trajectory matching. Zero stored keys. Fork the reference impl and stop shipping blind. 🧵👇</p>
<p>Threads/Bluesky:</p>
<p>The first managed AI agent regression gate that lives in your existing CI/CD. No new dashboard. No new vendor. Just: PR opens → agent deploys → behavioural tests run → scores &lt; 0.8 = build fails. This is how you ship agents responsibly. 🤖✅</p>
<p>Originally published on [your blog]. Republished with permission.</p>
<p>write according to hasnode community guidelines</p>
<p>title: "AWS Just Put AI Agent Regression Testing in GitHub Actions — Your PRs Now Have a Quality Gate" subtitle: "Bedrock AgentCore Evaluations + GitHub Actions = automated behavioural tests on every push. If your agent regresses, the build fails." slug: "aws-ai-agent-regression-testing-github-actions-bedrock-agentcore" tags: ["aws", "ai-agents", "github-actions", "ci-cd", "bedrock", "agentcore", "regression-testing", "devops"] cover: "photo-1558494949-ef010cbdcc31 (1600×900)" canonical: "" publish: true The problem: You changed your agent's system prompt. Or swapped the model. Or tweaked a tool description. The deployment succeeded — but the agent now hallucinates on refund requests. You find out when a customer complains.</p>
<p>The fix: AWS just dropped a reference implementation that runs automated agent evaluations inside GitHub Actions. Every PR gets a behavioural test suite. Scores drop below your threshold? The build fails. The merge stays blocked.</p>
<p>What Actually Shipped Not a blog post — a working reference pipeline you can fork today:</p>
<p>Strands-based agent on Amazon Bedrock AgentCore Runtime MCP server for tool exposure (with role-based access control) CDK infrastructure-as-code — runtime, Cognito (M2M + user-scoped auth), IAM roles GitHub Actions workflow using OIDC federation (no long-lived AWS keys) AgentCore Evaluations API — four built-in evaluators, threshold at 0.8 CloudWatch dashboards for trend monitoring The workflow: PR opens → CDK deploys dev stack → AgentCore spins up agent → evaluation prompts fire → scores computed → quality gate passes or fails the job.</p>
<p>The Four Evaluators That Caught the Regression In AWS's deliberate test, they broke the system prompt on purpose. Here's what flagged it:</p>
<p>Evaluator What It Measures Result on Broken Prompt GoalSuccessRate Did the agent actually complete the task? ❌ Failed Correctness Is the answer factually right? (LLM-as-a-judge) ❌ Failed ToolParameterAccuracy Did it pass the right arguments to tools? ❌ Failed ToolSelectionAccuracy Did it pick the correct tool at all? ✅ Passed Key insight: The agent still knew which tool to call — but couldn't use it properly. A single aggregate score would've masked the regression. Granular evaluators &gt; one number.</p>
<p>How the Quality Gate Works</p>
<h1>Simplified from the reference workflow</h1>
<ul>
<li>name: Evaluate Agent run: | python evaluate.py<br />--agent-id ${{ steps.deploy.outputs.agent-id }}<br />--threshold 0.8<br />--evaluators GoalSuccessRate,Correctness,ToolParameterAccuracy,ToolSelectionAccuracy Threshold: 0.8 / 1.0 (configurable) Failure mode: Script exits non-zero → GitHub Actions job fails → PR blocked Recovery: Fix the regression → push → re-eval → green checkmark Why This Matters: Agents Fail Differently Traditional software: passes tests today → passes same tests tomorrow. AI agents: handled refund perfectly yesterday → mishandles same request today because:</li>
</ul>
<p>Prompt was tweaked Model upgraded Tool description reworded LLM non-determinism (same input, different output) No stack trace. No clear error. Just... wrong behaviour.</p>
<p>AgentCore Evaluations (GA since March 2026, preview at re:Invent 2025) adds the measurement layer agents need:</p>
<p>Online: continuous eval on sampled production traffic On-demand: exact interactions you choose Batch: async scoring against CloudWatch Logs Built-in evaluators cover: Helpfulness, Correctness, GoalSuccessRate, ToolSelectionAccuracy, ToolParameterAccuracy, Safety, and more. Custom evaluators via LLM-as-a-judge or Lambda-based code.</p>
<p>The Auth Architecture Worth Stealing Three-layer pattern, zero stored credentials:</p>
<p>GitHub OIDC token → AWS validates → temporary IAM credentials (scoped to run) Cognito pool serves both M2M (agent→MCP) and user-scoped flows MCP server enforces role-based access control on tools This is the keyless CI/CD pattern AWS has been pushing — and it's production-ready here.</p>
<p>Trajectory Matching: The Hidden Gem Beyond LLM-as-a-judge, AgentCore does trajectory matching — comparing recorded tool calls against ground truth:</p>
<p>Exact-order: calls must match sequence precisely In-order: calls appear in order, but extras allowed Any-order: calls present, sequence doesn't matter This catches silent regressions where the agent picks the right tools but in the wrong order — or calls extra tools it shouldn't.</p>
<p>Competitors &amp; Context Tool Approach Gap LangSmith Tracing + eval Not AWS-native, separate platform Braintrust Eval framework Requires own infra AgentCore Evaluations Managed, integrated, CI/CD-native AWS-only (by design) If you're on AWS, this is the first managed evaluation layer that plugs straight into your existing CI/CD. No new dashboard, no new vendor, no credential sprawl.</p>
<p>What to Actually Do With This Fork the repo: aws-samples/agentcore-evaluations-github-actions (search it) Swap your agent into the Strands + MCP template Define your evaluation dataset — representative prompts + expected behaviours Set your threshold — start at 0.8, tighten as you build confidence Wire it to main branch protection — required status check = "Agent Evaluation" The Bigger Shift: Eval as Infrastructure This isn't a testing tool. It's infrastructure.</p>
<p>Build → CDK Deploy → AgentCore Runtime Observe → AgentCore Observability (traces, CloudWatch) Evaluate → AgentCore Evaluations (this release) Gate → GitHub Actions quality gate The loop closes. Every change gets measured before it ships.</p>
<p>Bottom Line AI agents in production without regression testing is shipping blind. AWS just gave you the reference implementation to stop doing that — managed evaluation, keyless CI/CD, granular scores, trajectory matching, all inside GitHub Actions you already use.</p>
<p>If you run agents on Bedrock: fork this today. The cost of a regression in production (wrong refund, leaked data, compliance violation) dwarfs the hour it takes to wire this up.</p>
<p>📚 References &amp; Further Reading AWS Machine Learning Blog:</p>
<p><a href="https://aws.amazon.com/blogs/machine-learning/automated-agent-evaluation-with-amazon-bedrock-agentcore-and-github-actions/">Automated agent evaluation with Amazon Bedrock AgentCore and GitHub Actions Amazon Bedrock AgentCore Evaluations Documentation</a></p>
<p>All code patterns and architectural descriptions reflect the reference implementation as of September 2026.</p>
]]></content:encoded></item><item><title><![CDATA[Chrome Zero-Day CVE-2026-85046 Is Being Exploited — Update Now]]></title><description><![CDATA[That “Relaunch Chrome” button you've been ignoring?
Maybe don't ignore it this time. 😬
Google has patched a high-severity Chrome zero-day, CVE-2026-85046, after confirming that attackers are already ]]></description><link>https://vikramkumarshukla.hashnode.dev/chrome-zero-day-cve-2026-85046-is-being-exploited-update-now</link><guid isPermaLink="true">https://vikramkumarshukla.hashnode.dev/chrome-zero-day-cve-2026-85046-is-being-exploited-update-now</guid><category><![CDATA[cybersecurity]]></category><category><![CDATA[chromezeroday]]></category><category><![CDATA[Google]]></category><category><![CDATA[zero_day_vulnerability]]></category><dc:creator><![CDATA[Vikram Kumar  Shukla]]></dc:creator><pubDate>Fri, 04 Sep 2026 09:04:21 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a999c97bb108e64b49f8dfd/1e6c93ba-5e81-4e8e-9a1d-bfff0ef24e9e.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>That <strong>“Relaunch Chrome”</strong> button you've been ignoring?</p>
<p>Maybe don't ignore it this time. 😬</p>
<p>Google has patched a <strong>high-severity Chrome zero-day, CVE-2026-85046</strong>, after confirming that attackers are already exploiting it in the wild.</p>
<p>The bug lives inside <strong>V8</strong>, Chrome's JavaScript and WebAssembly engine.</p>
<p>And yes — a malicious webpage can potentially be enough to trigger the problem.</p>
<hr />
<h2>🧨 So, What's Actually Broken?</h2>
<p><strong>CVE-2026-85046</strong> is a <strong>type confusion vulnerability</strong> in V8 with a <strong>CVSS score of 8.8</strong>.</p>
<p>Think of type confusion like this:</p>
<blockquote>
<p>You tell the browser, “This is a box of apples.”</p>
</blockquote>
<p>The browser trusts you.</p>
<p>Except someone secretly swapped the apples for USB drives. 🍎➡️💾</p>
<p>Now the browser starts handling the object using assumptions that aren't true.</p>
<p>In V8, that kind of mistake can become much more serious.</p>
<p>Security researcher <strong>Salvatore Gulizia (@Serotav)</strong> discovered the vulnerability and reported it to Google on <strong>August 4, 2026</strong>.</p>
<p>His research showed that the bug could potentially be turned into <strong>arbitrary read/write access to the JavaScript heap</strong> — a much more dangerous primitive than a simple browser crash.</p>
<hr />
<h2>☠️ The Part That Matters Most</h2>
<p>Here's the big red flag:</p>
<p><strong>Google says an exploit for CVE-2026-85046 already exists in the wild.</strong></p>
<p>Google hasn't publicly revealed the full attack details yet.</p>
<p>That's deliberate.</p>
<p>Publishing every exploitation detail before most users patch would basically hand attackers a better instruction manual.</p>
<p>So for now, we know the vulnerability is being exploited — but not exactly <strong>who is behind the attacks or who they're targeting</strong>.</p>
<hr />
<h2>🤯 Fun Fact: Your Browser Is Basically an Operating System Now</h2>
<p>Modern browsers aren't just for opening websites.</p>
<p>Chrome runs:</p>
<ul>
<li><p>JavaScript</p>
</li>
<li><p>WebAssembly</p>
</li>
<li><p>Web apps</p>
</li>
<li><p>Games</p>
</li>
<li><p>Video applications</p>
</li>
<li><p>Cloud software</p>
</li>
<li><p>Authentication flows</p>
</li>
<li><p>Password managers</p>
</li>
<li><p>Enterprise applications</p>
</li>
</ul>
<p>That's why browser vulnerabilities are such valuable targets.</p>
<p>A compromised browser can potentially become the starting point for a much bigger attack.</p>
<p><strong>Your browser is part of your security perimeter.</strong></p>
<hr />
<h2>📊 And This Isn't Chrome's First Zero-Day This Year</h2>
<p>CVE-2026-85046 is reportedly the <strong>sixth actively exploited Chrome zero-day patched by Google in 2026</strong>.</p>
<p>Previous vulnerabilities include:</p>
<ul>
<li><p>CVE-2026-2441</p>
</li>
<li><p>CVE-2026-3909</p>
</li>
<li><p>CVE-2026-3910</p>
</li>
<li><p>CVE-2026-5281</p>
</li>
<li><p>CVE-2026-11645</p>
</li>
</ul>
<p>That doesn't mean Chrome is uniquely insecure.</p>
<p>It means browsers are <strong>high-value targets</strong>.</p>
<p>Billions of users + massive attack surface = very attractive target.</p>
<hr />
<h2>🔧 What Should You Do?</h2>
<p>Pretty simple.</p>
<p>Open:</p>
<p><strong>Chrome → ⋮ → Help → About Google Chrome</strong></p>
<p>Check for the latest update and hit <strong>Relaunch</strong> when prompted.</p>
<p>Google's patched desktop versions are:</p>
<p><strong>Windows/macOS:</strong> <code>152.0.7977.82/.83</code><br /><strong>Linux:</strong> <code>152.0.7977.82</code></p>
<p>If you're using another Chromium-based browser such as Edge, Brave, Opera or Vivaldi, check for its corresponding security update too.</p>
<hr />
<h2>🧠 One Last Thing</h2>
<p>Zero-days sound like something that only matters to security researchers.</p>
<p>They don't.</p>
<p>The dangerous part of a browser zero-day is that the victim may not realize anything unusual happened.</p>
<p>No scary terminal.</p>
<p>No obvious malware popup.</p>
<p>Sometimes, just <strong>a browser tab</strong>.</p>
<p>So if Chrome is asking for a restart today, don't postpone it for “later.”</p>
<p><strong>Patch now. Scroll later.</strong> 🔒</p>
<hr />
<h3>🔗 Source</h3>
<p>Based on Google's Chrome security release and public technical research from Salvatore Gulizia regarding CVE-2026-85046.</p>
<p><em>This article is an original editorial rewrite and analysis, not a reproduction of the source article.</em></p>
]]></content:encoded></item></channel></rss>