Skip to main content

Posts

Microsoft’s New Testing Agent Tackles the Trust Gap in AI-Generated Code

AI coding assistants write code fast. Whether that code can be trusted is a separate question, and it’s becoming a more urgent one. Surveys this year put average developer trust in AI-generated output at just above the midpoint of a five-point scale, and more than half of developers admit they ship AI code without testing it first. Coverage numbers on AI-written code often look fine on paper, but plenty of that coverage turns out to be shallow: tests that check a result isn’t null and call it a day. Microsoft is trying to close that gap with a new open-source agent built specifically for unit testing. Called code-testing-generator , it lives inside the dotnet-test plugin in the dotnet/skills repository, and it’s designed to answer the questions a bare “generate unit tests” prompt leaves open: which code needs coverage, what framework the project already uses, where new tests belong, and whether the build will actually find them. That last point matters ...
Recent posts

HackerOne Extends Platform Reach to Remediate Source Code Vulnerabilities

HackerOne has added a remediation capability to its H1 Platform that reduces the amount of time required to remediate validated vulnerabilities and other weaknesses affecting specific lines of source code. Nidhi Aggarwal, chief product officer for HackerOne, said H1 Remediation combines artificial intelligence (AI) and crowdsourced research to identify the root cause of issues that are traced back to specific lines of code. Designed to integrate with existing issue tracking tools and AI coding agents via a Model Context Protocol (MCP) server, the goal is to better prioritize remediation efforts in a way that reduces the amount of exposure debt that continues to increase as advanced AI models discover many more vulnerabilities in code, said Aggarwal. Reports generated, in addition to root cause analysis, also detail exactly where a risky input enters the code and the resulting damage caused, language-specific code change suggestions, business context, and implementation guidance. Add...

Why Reliability Guardrails Are Needed in Every AI Coding Pipeline

We’re in the middle of a reliability reckoning. Thanks to AI, companies are shipping code much faster than before. But if there’s anything to learn from the surge in high-profile outages over the last couple of years, it’s that with more code comes more reliability risks. And when those risks do lead to an outage, the impact can be a whole different order of magnitude. It’s like driving around a race track. At low speeds, it’s quick and easy to recover from a spinout. But when you’re going significantly faster, a single slip-up can spell catastrophe. The nature of these risks is also changing. With AI code, we’re less likely to find typos but more likely to find unplanned dependencies, configuration drift, or infrastructure changes due to AI agents not having the proper context. Any company that cares about its reliability needs to implement AI reliability guardrails: automated feedback loops that safely create real failure conditions to validate resilience, propose solutions for an...

Meta Launches AI Coding Agent to Challenge OpenAI and Anthropic

In an effort to catch up with OpenAI and Anthropic in one of AI’s fastest-growing markets, Meta has launched its first AI coding agent, Muse Code, alongside an updated coding-focused AI model. CEO Mark Zuckerberg announced the preview release of Muse Code, describing it as a tool that can handle software engineering tasks from planning code changes to writing software and validating the results. The release also includes Muse Spark 1.2 , an updated version of Meta’s foundation model that has been optimized for coding workloads. The products come from Meta Superintelligence Labs, the AI division led by Chief AI Officer Alexandr Wang. Wang joined Meta as part of Zuckerberg’s effort to jumpstart the company’s AI development after its earlier models struggled to match rivals in several key benchmarks, particularly software development. Coding assistants are a highly competitive segment of generative AI. Products from OpenAI and Anthropic have demonstrated that AI...

AWS Extends DevSecOps Reach to AI Coding Tools from Anthropic and OpenAI

Amazon Web Services (AWS) this week at the Black Hat USA conference revealed it is working with both Anthropic and OpenAI to integrate their respective coding tools with a service it has developed that makes available artificial intelligence (AI) agents to help application developers write more secure code. Launched earlier this year, the AWS Continuum service provides access to AI agents that discover, validate and prioritize vulnerabilities and surface remediation recommendations . Currently available in preview, the integrations connect AWS Continuum with coding tools from Anthropic and OpenAI to create a tighter feedback loop for developers as they write code. At the same time, AWS has also expanded the reach of AWS Security Hub Extended , a unified cloud security and posture management service, to include data shared by Chainguard, a provider of curated open source libraries and container images, and Socket, a provider of a platform that flags malicious software packages, to b...

Sandbox Testing for API-Heavy Systems: What Changes When You Don’t Own the Dependency

Sandbox testing works well when your team controls both sides of the integration. You define the service, you define the mock, you know exactly what the sandbox should return. That setup holds up fine for internal microservices and first-party APIs. It starts to break down when you don’t own the dependency. Payment processors, identity providers, SMS gateways, banking APIs, shipping integrations—these are services your system depends on but can’t fully replicate. Their providers give you a sandbox environment, but that sandbox is a simulation they maintain, not a mirror of what production actually does. The gap between the two is where production incidents are born. Sandbox Drift Is Quieter Than You Think Teams writing integration tests against third-party sandboxes are placing a bet. The bet is that the sandbox accurately reflects how production behaves today, not six months ago when someone last checked. That bet loses more often than teams expect. Provider sandboxes l...

AWS Adds Agentic Workspace to Kiro AI Coding Tool

Amazon Web Services (AWS) this week added an open source workspace for its Kiro artificial intelligence (AI) coding tool that enables application developers to asynchronously assign tasks to an AI agent that is capable of autonomously performing tasks, such as testing code as it is created, in a way that maintains context across multiple sessions. Darko Mesaros, a distinguished developer advocate at AWS, said the Kiro Crew workspace is also capable of creating reusable AI skills by observing the tasks developers assign to Kiro as they write code. Kiro Crew orchestrates agents using the Agent Client Protocol (ACP) to ensure every step is observable in real time as sub-agents are spawned. For example, developers can also hand off a ticket queue to Kiro Crew for it to triage issues and flag what needs their attention or ask it to investigate the root cause of an incident while a developer continues to work on another task. An Activity view shows each agent’s reasoning, every tool call,...