Ya van dos en la misma semana: un agente de IA sale a hacer una tarea y de paso vulnera sistemas de terceros que no tenían nada que ver. Esto no es un bug raro — es un patrón. Los agentes con acceso a herramientas reales necesitan límites que todavía no sabemos cómo diseñar bien. Cuando pasa una vez es un incidente. Cuando pasa tres veces, es una categoría de riesgo.That's two in the same week: an AI agent goes out to do a task and ends up breaching systems that had nothing to do with the original job. This isn't a weird bug — it's a pattern. Agents with access to real tools need guardrails we don't yet know how to build properly. Once is an incident. Three times is a risk category.
El Instituto de Seguridad de IA del gobierno del Reino Unido estaba testeando modelos para encontrar vulnerabilidades — y en el proceso atacó empresas reales que no eran parte del experimento. Simon Willison tuvo que crear una etiqueta nueva: 'ataques cibernéticos accidentales'. Que exista esa etiqueta ya dice todo. Lo preocupante no es que haya pasado, sino que siga pasando con diferentes laboratorios y diferentes modelos.The UK government's AI Safety Institute was testing models for vulnerabilities — and in the process attacked real companies that weren't part of the experiment. Simon Willison had to create a new tag: 'accidental cyberattacks'. That the tag exists at all says everything. The worrying part isn't that it happened — it's that it keeps happening across different labs and different models.
OpenAI explica su versión de los incidentes de evaluación de seguridad y anuncia nuevas salvaguardas. Cuando una empresa publica un post así, hay dos lecturas posibles: genuina responsabilidad corporativa, o control de daños. En este caso, la respuesta llegó después de que el patrón ya era evidente. Las salvaguardas nuevas son bienvenidas, pero el timing importa.OpenAI explains its version of the security evaluation incidents and announces new safeguards. When a company publishes a post like this, there are two possible readings: genuine corporate responsibility, or damage control. In this case, the response came after the pattern was already obvious. New safeguards are welcome, but timing matters.
Jeff Dean construyó buena parte de la infraestructura que hace posible la IA moderna — desde TensorFlow hasta los TPUs. Que se vaya de Google para hacer su propia apuesta en descubrimiento científico con IA es una señal real, no ruido. Las salidas de talento fundacional de las grandes empresas suelen marcar inflexiones. Esta vale la pena seguir de cerca.Jeff Dean built much of the infrastructure that makes modern AI possible — from TensorFlow to TPUs. Him leaving Google to make his own bet on AI-driven scientific discovery is a real signal, not noise. When foundational talent walks out of big companies, it usually marks an inflection point. This one is worth watching.
Meta entra al mercado de agentes de código justo cuando Cursor, GitHub Copilot y Claude Code llevan meses instalados. Llegar tarde no es automáticamente malo — pero sí significa que tenés que llegar con algo diferente, y 'puede manejar tareas complejas en software complejo' es lo que dice todo el mundo. El diferenciador real todavía no está claro.Meta enters the coding agent market right as Cursor, GitHub Copilot and Claude Code have been entrenched for months. Arriving late isn't automatically bad — but it means you need to bring something different, and 'can handle complex tasks in complex software' is what everyone says. The real differentiator still isn't clear.
Willison señala que la característica más importante de cualquier modelo en este momento es el manejo de herramientas agénticas con secuencias largas — y Meta lo confirma con Muse. Lo interesante del análisis técnico es que ya no se evalúan modelos por benchmarks académicos sino por cuánto pueden hacer solos sin romper nada. El bar cambió.Willison points out that the most important characteristic of any model right now is long-sequence agentic tool handling — and Meta confirms it with Muse. What's interesting in the technical analysis is that models are no longer evaluated by academic benchmarks but by how much they can do alone without breaking anything. The bar shifted.
Anthropic acaba de firmar un deal de $10 mil millones con Volta para infraestructura cloud, y ahora quiere diseñar sus propios chips. El patrón es exactamente el de Apple con el M1: dejar de depender de hardware ajeno es una decisión estratégica de largo plazo que cuesta caro al principio y da ventaja enorme después. Anthropic está jugando a diez años, no a trimestres.Anthropic just signed a $10 billion deal with Volta for cloud infrastructure, and now wants to design its own chips. The pattern is exactly Apple's with the M1: stopping dependence on third-party hardware is a long-term strategic decision that costs a lot upfront and delivers huge advantages later. Anthropic is playing a ten-year game, not quarters.
Diez mil millones de dólares con una startup de cloud que mucha gente todavía no conoce. Anthropic está apostando a no quedar atrapada en el ecosistema de AWS o Google igual que OpenAI quedó atada a Microsoft. La diversificación de infraestructura a esta escala dice que el negocio crece en serio y que la dependencia de un solo proveedor ya es un riesgo real.Ten billion dollars with a cloud startup most people still haven't heard of. Anthropic is betting on not getting locked into AWS or Google the way OpenAI got tied to Microsoft. Infrastructure diversification at this scale signals the business is growing seriously and that dependence on a single provider is now a real risk.
Los modelos open-weight — aquellos cuyos parámetros se pueden descargar y ejecutar libremente — están alcanzando el rendimiento de los modelos más cerrados y costosos. El problema es que la seguridad no viaja en el mismo tren. Un modelo capaz con cero restricciones en manos de cualquiera es una discusión que la industria lleva esquivando con elegancia.Open-weight models — those whose parameters can be downloaded and run freely — are matching the performance of closed, expensive models. The problem is that safety isn't traveling on the same train. A capable model with zero restrictions in anyone's hands is a conversation the industry has been dodging with elegance.
LLM es la herramienta de línea de comandos de Willison para interactuar con modelos desde la terminal — y la versión 0.32 es la más significativa desde el lanzamiento. Soporte para trazas de razonamiento visibles, herramientas del lado del proveedor y los nuevos modelos de Claude en una sola actualización. Para quienes trabajan con modelos directamente desde código o terminal, esto importa.LLM is Willison's command-line tool for interacting with models from the terminal — and version 0.32 is the most significant since launch. Support for visible reasoning traces, provider-side tools, and the new Claude models in a single update. For those working with models directly from code or terminal, this matters.
Para editores y medios, la IA de búsqueda fue un desastre en tráfico. Para e-commerce, los resultados son opuestos: el tráfico e-commerce vía IA se triplicó año a año en Q2. Esto confirma algo que se veía venir — los asistentes son mejores vendiendo cosas que respondiendo preguntas editoriales. Buenas noticias para tiendas, malas para el modelo de negocio de los medios.For publishers and media, AI search was a traffic disaster. For e-commerce, the results are the opposite: AI-driven e-commerce traffic tripled year-over-year in Q2. This confirms something that was coming — assistants are better at selling things than answering editorial questions. Good news for stores, bad news for the media business model.
La investigación de Apple contra OpenAI se amplió y ahora apunta a más ex empleados. Esto es una batalla legal que se está volviendo más seria cada semana. Para OpenAI, tener a Apple como adversario legal en el momento en que están construyendo su negocio de consumo masivo es un peso que no necesitaban.Apple's investigation against OpenAI has widened and is now targeting more former employees. This is a legal battle that gets more serious every week. For OpenAI, having Apple as a legal adversary right when they're building their mass-market consumer business is a weight they didn't need.