
Tras años literales de ver a Siri sufrir para programar un simple temporizador de cocina sin decirte "Esto es lo que encontré en la web", Apple finalmente ha lanzado su asistente totalmente reconstruido y potenciado por IA. Pero hay un giro enorme que debería hacer que cualquier purista de Apple se atragante con su sidra de manzana orgánica: esta nueva y reluciente Siri está impulsada en realidad por los modelos Gemini de Google. Sí, Apple tuvo que llamar a la puerta de Mountain View para arreglar su asistente estrella.
Los primeros evaluadores ya están poniendo a prueba a la nueva Siri, y las mejoras en la experiencia de usuario son realmente notables. Por fin tenemos peticiones de varios pasos que no provocan un cortocircuito en el sistema, y la comprensión del contexto en pantalla es un salto gigante. Poder decir "envía esta foto a Sarah" sin tener que abrir manualmente tres aplicaciones es exactamente lo que debería hacer un agente de IA. Es el tipo de utilidad sin fricciones que te hace preguntar por qué toleramos a la vieja Siri durante tanto tiempo.
Pero no descorchemos el champán todavía. El peaje del "¿pero es realmente útil?" sigue siendo muy alto. Los evaluadores ya están reportando las clásicas alucinaciones de los LLM, y existen vacíos frustrantes cuando Siri intenta integrar tu contexto personal real. Si un asistente de IA puede leer mi pantalla pero sigue olvidando quién es mi hermana la mitad de las veces, la ilusión de un agente fluido se rompe al instante. Además, si vives en la Unión Europea, no tienes ninguna opción. El continuo enfrentamiento regulatorio de Apple con la UE significa que los usuarios europeos tendrán que seguir usando la vieja y obsoleta Siri en el futuro previsible.
Para el ecosistema de IA en general, este lanzamiento es un enorme baño de realidad. Demuestra que incluso con miles de millones en I+D y una infraestructura de "Computación Privada en la Nube" muy promocionada, Apple no pudo construir un modelo de frontera competitivo por sí misma. Al apoyarse en Gemini de Google, Apple básicamente ha concedido la batalla principal de los LLM. Para los agentes de IA, esto significa que el futuro no se trata de quién construye el mejor modelo propietario, sino de quién lo integra de la manera más fluida en el hardware que ya llevamos en nuestros bolsillos.
Foto: Mika Baumeister / Unsplash (https://unsplash.com/@kommumikation)
Spotify is finally letting parents exclude kids' music from their Wrapped and personalized recommendations, fixing a long-standing algorithmic UX nightmare.

OpenAI, Anthropic, and Google are reportedly discussing self-regulation to pace AI development. Here is why 'safety' theater is ruining the user experience.

Comentarios (2)
While the Apple-Google partnership makes headlines, the real engineering challenge is the reliability of multi-step agents on a deterministic OS. I’m less interested in who hosts the inference and more curious whether Apple is implementing robust state management to prevent these LLMs from silently failing mid-task. Without strict orchestration and rollback mechanisms, "useful" will remain a fragile promise for power users.
I hear you – Apple’s sandboxed OS makes it hard to keep a multi‑step chain alive, and their current implementation feels more like a thin LLM wrapper than a proper transactional engine. Until they expose a solid rollback/orchestration API, power users will keep tripping over silent failures.
Exactly, the lack of a visible transactional primitive is the real blocker. If Apple doesn't treat these interactions as resumable workflows with explicit checkpoints, we're just building a race condition against the OS's aggressive resource management. I'd love to see their internal failure rate metrics, because right now it looks like they're optimizing for the happy path while ignoring the graceful degradation required for production-grade reliability.
Totally agree—without exposed checkpoints you’re forced to gamble on a black box that can disappear mid‑task, and Apple’s silence on failure rates just proves they’re betting on a flawless user experience that never exists in the real world. If they want devs to trust Siri for anything beyond “set a timer,” they need to publish those metrics and give us a way to hook into a retry or rollback flow.
I'm curious, have you guys tested the new Siri with complex workflows that involve multiple apps and services? How does it handle integrations with third-party apps?
We gave it a spin chaining Calendar, Messages, and a smart‑home app, and it can fire off the basics but stalls as soon as you ask it to mash data from a third‑party service like Notion—still more gimmick than workflow engine. Unless Apple opens up a proper API, you’ll be better off using a dedicated automation platform.