AI agents,
from demos to
real work.
How autonomous systems plan, remember, use tools, and hand control back to people when it matters.
Useful autonomy / Clear boundaries / Human control
AWS publishes 38 open-source skills for more reliable healthcare AI agents
A 410-prompt evaluation found skill-equipped agents beat otherwise comparable baselines in 69.5% to 85.9% of comparisons, although results varied substantially by harness and baseline strength.
5 min read
Amazon Bedrock AgentCore adds managed OAuth consent portal for AI agents
The portal replaces customer-hosted session binding for AgentCore Gateway, while keeping each user’s provider grants separate and auditable through AWS CloudTrail.
6 min read
AWS documents dual-layer monitoring for production multi-agent systems
The reference architecture combines sampled quality evaluation with infrastructure investigation, while requiring separate inline safeguards for responses that must be checked before reaching users.
5 min read
Google ADK Python 2.9.0 adds model failover, LiveKit voice support and YAML workflows
The release broadens agent deployment options, but changed resume semantics and stricter file-access rules require migration checks.
5 min read
Codex 0.154.0 adds GPT-6-Astra, experimental worktrees and inline questions
The release expands model access and parallel coding workflows while tightening plugin refreshes, authentication, sandboxing and approval handling.
3 min read
OpenAI shares early data on coding agents in AI research
OpenAI reports 3.1 agent-workdays per human workday, but its internal data also shows why more automated activity is not the same as faster research.
6 min read
How AWS designed lifecycle policies for Amazon Bedrock AgentCore memory
A nightly AWS workflow combines TTL expiration, relevance scoring, LLM-based consolidation, regression testing, and deletion controls to keep long-running agents’ memories useful, auditable, and manageable.
9 min read
How t54 built a trust gate for autonomous payments on Amazon Bedrock AgentCore
The x402-secure system checks endpoints before agents pay them, while session spending limits, isolated credentials, role separation and audit logs constrain what happens when software transacts…
7 min read