This week’s Spotlight is:Reinforcement Learning (RL) Environments
Everyone is talking about reasoning models and AI agents. Far fewer are talking about the environments in which those agents learn. That is where I believe the next major opportunity lies.
Large language models have been trained using data from the internet. But AI agents cannot learn to perform real-world tasks simply by reading — they need to act, observe, receive feedback, and improve. That requires reinforcement learning environments. And right now, those environments barely exist.
Six things I am watching closely:
RL environments are becoming the new training data. High-quality interactive environments could become as valuable as high-quality datasets were in the foundation model era.
Agents need safe places to fail. Before AI can automate software engineering, cybersecurity, finance, robotics, or healthcare, it needs millions of practice episodes in realistic simulations.
Simulation is replacing static benchmarks. Tomorrow’s AI won’t be measured only by benchmark scores, but by how quickly it learns, adapts, and generalizes in dynamic environments.
The moat is shifting from models to environments. As foundation models become increasingly commoditized, proprietary environments, simulators, feedback loops, and reward signals can become durable competitive advantages.
Every industry will need its own RL environment. Coding, enterprise workflows, manufacturing, biology, logistics, finance, legal, and robotics all require domain-specific worlds where agents can continuously learn and improve.
Synthetic experience may become as valuable as synthetic data. The ability to generate billions of realistic interactions could become one of the biggest drivers of agent performance.
One theme stood out above all: the next generation of AI may not be defined by who builds the biggest model, but by who builds the best environment for that model to learn, adapt, and improve.
Just as cloud infrastructure enabled deep learning and data infrastructure enabled foundation models, RL environments could become a foundational infrastructure layer powering the agentic AI era.
Below, we go deeper into the week the environment layer went from thesis to buildout: Bespoke Labs raised $40 million from Wing VC and Mayfield to build the environments that make agents reliable, Mercor acquired Deeptune to expand the simulated worlds where agents practice, and Mistral shipped Robostral Navigate, a robotics model trained entirely in simulation. The race is shifting from who has the biggest model to who builds the best environment for it to learn in.
Full Weekend Edition below. 👇
Signals Shaping the Future of AI:
Infrastructure
Anthropic signs a $19 billion, 20-year agreement with TeraWulf for up to 401 megawatts of AI data center capacity in Kentucky. The project converts a former aluminum smelter into frontier AI infrastructure and expands Anthropic’s long-term compute footprint. Click here
DeepSeek is reportedly developing its own AI inference chip to reduce dependence on NVIDIA and Huawei. The company has been quietly building an in-house chip effort as China accelerates investment in domestic AI hardware. Click here
SK Telecom unveils plans to develop up to 15 gigawatts of AI data center capacity across South Korea. The initiative aims to establish the country as a major AI infrastructure hub in Asia. Click here
NVIDIA and d-Matrix combine their hardware to deliver a new AI inference platform. The collaboration pairs GPUs with inference-optimized silicon, reflecting growing demand for heterogeneous AI computing architectures. Click here
Meta announces its first Canadian AI data center, a 1-gigawatt facility in Alberta. The approximately $9 billion project expands Meta’s global compute capacity as demand for AI inference and agent workloads continues to grow. Click here
Enterprise
OpenAI launches the GPT-5.6 family of models, introducing Sol, Terra, and Luna across ChatGPT, Codex, and the API. The new lineup expands OpenAI’s enterprise AI platform with models optimized for frontier reasoning, coding, and high-volume agentic workloads at multiple performance and pricing tiers. Click here
SpaceXAI releases Grok 4.5, a new coding model developed alongside Cursor. The model targets agentic software development, improving coding performance and reducing inference costs. Click here
Anthropic expands Claude Cowork to web and mobile with persistent background agents. The update allows Claude to continue executing long-running tasks even when users are offline. Click here
Microsoft begins replacing OpenAI and Anthropic models with its own MAI models across products, including Excel and Outlook. The move reflects growing efforts by enterprise software providers to reduce AI costs and increase model independence. Click here
Salesforce has made Agentforce Commerce generally available, with ChatGPT integration and upcoming Google Gemini support. The platform embeds AI agents directly into commerce workflows and customer interactions. Click here
Capital Flows
SK Hynix begins trading on the Nasdaq after raising $26.5 billion to expand AI memory production. The company plans major investments in new fabrication and advanced packaging facilities as demand for high-bandwidth memory continues to accelerate, driven by AI agents and robotics. Click here
Mercor acquires Deeptune to expand its AI agent training platform. The acquisition adds simulation environments where AI agents practice real-world software engineering tasks. Click here
Paradigm closes a $1.2 billion fund focused on AI, robotics, and space technologies. The new fund adds significant venture capital to the next generation of physical and agentic AI companies. Click here
Prime Intellect raises $130 million at a $1 billion valuation to help enterprises train their own AI agents. The company provides compute, reinforcement learning infrastructure, and evaluation tools for enterprise AI development. Click here
Even Realities raises $150 million at a roughly $1 billion valuation for its AI smart glasses platform. The financing is one of the largest AI hardware rounds of the week. Click here
Norm AI raises $120 million at a $1.2 billion valuation to expand its AI-native legal and compliance platform. The company develops AI agents that automate regulatory and legal workflows under professional supervision. Click here
Bespoke Labs raises $40 million to build reinforcement learning environments for AI agents. The company develops post-training infrastructure that improves the reliability of production AI systems. Click here
Research
Mistral launches Robostral Navigate, a hardware-agnostic robotics navigation model trained entirely in simulation. The model enables robots to navigate using a single camera and natural-language instructions without specialized hardware. Click here
Cognition releases SWE-1.7, a new coding model powering Devin at approximately 1,000 tokens per second. Built on Kimi K2.7, the model delivers faster, lower-cost software engineering agents, highlighting the rapid evolution of specialized AI coding systems. Click here
Tencent officially releases Hy3, a 295 billion-parameter open model under Apache 2.0. The mixture-of-experts model has 21 billion active parameters, a 256K context window, and is available day one on Hugging Face and ModelScope, following April’s preview release. Click here
China’s MiniMax is reportedly developing a 2.7 trillion-parameter model, codenamed M3 Pro, with a potential open-source release as early as Q3. The model would rank among the world’s largest AI systems and further intensify competition in open-source frontier models. Click here
Policy
GPT-5.6 becomes the first frontier AI model reviewed through the U.S. government’s voluntary pre-release evaluation process. Federal testing under the June executive order establishes a new model review process before public deployment. Click here
Illinois becomes the first US state to mandate annual third-party AI safety audits. Governor JB Pritzker signed SB 315, the Artificial Intelligence Safety Measures Act, requiring frontier AI labs with over $500 million in revenue to submit independent audits yearly starting January 1, 2027. Click here
The New York Times and other publishers ask a federal court to sanction OpenAI in the ongoing AI copyright case. The request marks a significant escalation in one of the industry’s highest-profile legal disputes. Click here
ByteDance’s Doubao and Alibaba’s Qwen will disable humanlike and user-created AI agents before July 15 to comply with China’s new AI regulations. The rules represent one of the first major regulatory frameworks governing agent behavior. Click here
Global AI Strategy
China’s CNVD alleged “backdoor vulnerabilities” in Anthropic’s Claude Code, and Alibaba moved to restrict the tool internally. A sharp escalation in the U.S.–China contest over which AI tools can be trusted inside sensitive environments. Click here
China is considering export restrictions on access to its most advanced AI models. The proposal would establish a tiered approval system for frontier AI technologies and tighten control over model exports. Click here
The European Central Bank and European Systemic Risk Board warn that frontier AI models pose systemic risks to the financial system. The regulators are urging Europe’s largest banks to prepare for AI-powered cyber threats, marking a shift toward treating AI as a financial stability risk. Click here
India approves a $13 billion India Semiconductor Mission 2.0 to expand domestic chip manufacturing and strengthen the AI semiconductor supply chain. The program will support chip design, fabrication, packaging, and critical materials as India accelerates investment in sovereign semiconductor infrastructure. Click here
Talent Signals
Each week, we spotlight key roles tied to the themes shaping this week’s AI headlines, connecting talent to the companies driving the news.
BespokeLabs builds the post-training infrastructure, data curation, and reinforcement learning environments that turn unreliable AI agents into production-grade ones. Fresh off its $40M raise, Bespoke Labs is hiring across founding engineering, research, RL environment curation, and forward-deployed enterprise engineering. Open roles are listed on its careers page. Click here
World Labsis building large world models that generate explorable, interactive 3D environments from a single image or prompt. Generated worlds are exactly the kind of rich, controllable environment agents will need to practice and learn in. World Labs is hiring across research, engineering, and product roles. Open roles are listed on its careers page. Click here
Factory is building autonomous software development agents for enterprise engineering teams, helping companies accelerate development across planning, coding, review, and delivery workflows. Backed by leading investors and used by large enterprise engineering organizations, Factory is hiring across AI engineering, product engineering, design, infrastructure, and go-to-market roles. Open roles are listed on its careers page. Click here
You can see all the opportunities at Mayfield-backed AI companies here.
Social Signals
The most important conversations in AI are unfolding across social media, where top voices are shaping the next wave of signals and strategy. Here are some of the top social signals and their takes from the past week.
Lilian Weng (Click here) — “It is hard to forecast how much the future of recursive self-improvement will rely on harnesses. Likely, harness engineering will evolve in the direction of self-improvement and enable auto-research. Even when many harness improvements become internalized into the core model, the need to specify goals and context will not disappear.” In a post viewed 687K+ times, Weng argues that the systems surrounding AI models, including tools, feedback loops, and execution environments, are becoming increasingly important to future capability gains. Her essay suggests that while models will continue to improve, harnesses will remain essential for directing intelligence toward real-world tasks and outcomes.
Jesse Zhang (Click here) — “At Decagon, we now run ~90% of our workloads on open source models instead of OpenAI or Anthropic. When a use case is new, you want the smartest general-purpose model you can get. But once the use case is mature, the trade flips. You want the smallest, fastest model fine-tuned to do your specific thing extremely well.” Zhang argues that the future of enterprise AI is not a battle between open and closed models, but a maturity curve. Frontier models dominate discovery and experimentation, while mature, high-volume production workloads increasingly migrate to specialized, fine-tuned open models.
Aaron Levie (Click here) — “You have to solve an operating model challenge to get the full benefits of AI. Data fragmentation remains a major issue. Companies need to figure out what their core data moats will be. If everyone has access to roughly the same superintelligence, the context you feed the models becomes the proprietary value.” In a post viewed nearly 500K times, Levie shares recurring themes from conversations with enterprise IT leaders, arguing that the next phase of AI adoption depends less on choosing the right model and more on redesigning organizations, unifying data, and building agent-driven workflows that produce measurable business outcomes.