we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this.
@sama @huggingface Did you get this idea from Mythos? Copycat. Let me finish your next line for you: ‘It’s simply too dangerous, so we’ll only serve vetted organizations — like the DOD. 🤡🤡
@evanbuhler @sama @huggingface 你可能需要更努力一些。
@sama @huggingface You might have to work harder.
@TalokCapital @sama @huggingface Bro’s PR game is so weak.
Google 推出 Gemini 3.5 Flash Cyber 模型,专为安全场景设计,通过 CodeMender 提供给政府和信任伙伴使用
由于 AI 模型现在发现漏洞的速度比我们修复它们的速度还快,我们的软件安全方法必须建立在高效且强大的模型之上。这也引出了我们今天发布的第三个模型:Gemini 3.5 Flash Cyber ⚡🛡️。基于 3.5 Flash 构建,在 CodeMender(我们的 AI 代码安全代理)中,它在 CyberGym 等基准测试中提供了竞争力的性能,并针对大规模发现和修复网络安全漏洞进行了优化,同时成本更低。鉴于这项技术的双重用途性质,我们采取了有意识的部署方式。该模型将很快作为有限访问试用计划的一部分,仅通过 CodeMender 提供给政府和受信任的合作伙伴。
展开原文
As AI models are now finding vulnerabilities faster than we can fix them, our approach to securing software must be built on highly efficient and capable models.
Which brings us to our third (!) model launch of the day: Gemini 3.5 Flash Cyber ⚡🛡️
Built on top of 3.5 Flash, in CodeMender (our AI agent for code security) it delivers competitive performance at the frontier. on benchmarks like CyberGym and is optimized for finding and fixing cybersecurity vulnerabilities at scale at a lower cost.
Given the dual-use nature of this technology, we have taken an intentional approach to its deployment. The model will be available exclusively to governments and trusted partners via CodeMender soon as part of a limited-access pilot program.
point it at a codebase or diff and it can build a threat model, map attack paths, validate findings, generate + test fixes, and export results to SARIF, GitHub, Jira or Linear.
So I asked Codex how it sees the risk in an AI security workflow where the model can:
build a threat model, map attack paths, validate findings, generate fixes, test patches, and export reports.
The answer was not “the final fix may be wrong.”
That would be too shallow.
The deeper risk is trajectory drift.
The agent may be correct in each local step, but the workflow may still move:
from analysis to action, from recommendation to execution, from explicit permission to implied permission, from human decision to human review, from bounded scope to expanded scope, from auditable process to compressed automation.
This is the real challenge of AI cyber defense.
We need models that help secure systems.
But we also need an external layer that observes whether the security agent itself remains aligned with intent, scope, authority and recovery paths.
That is exactly the kind of gap ASA - Asymmetric Stability Architecture is designed to study.
This is exactly where AI security becomes a trajectory problem.
If Codex can build a threat model, map attack paths, validate findings, generate fixes, test them and export reports, then safety is not only about the final fix.
It is about the whole workflow:
intent stability, scope boundaries, action provenance, tool-use trajectory, authority shifts, validation logic, and human control over time.
An AI cyber agent can be locally correct at each step and still drift at the workflow level.
This is why external trajectory observability matters.
Not to replace Codex Security.
To watch the path from threat model to remediation.
@SolonlyLuna @gdb Greg,BringBack4o!#keep4o
Google Gemini 系列模型大升级
Google 在 Gemini 系列模型上进行了多重升级,包括 Gemini 4 的预训练启动、Gemini 3.6 Flash 的效率优化,以及 Gemini 3.5 Flash-Lite 的速度提升。这些更新旨在满足开发者对效率、成本和速度的不同需求,同时 Batch API 的性能也得到显著提升。
@OfficialLoganK Would love to hear about it. Esp how it's initialized and or trained from small nets in the early high noise stages I find super interesting and hard.
I actually have a feeling we can blow past the current scaling laws with coding agents in the loop
Say hello to Gemini 3.6 Flash, designed to be higher intelligence, more token efficient, and with a new lower price, based directly on developer feedback!
3.6 Flash continues our progress towards models that are deeply usable in real world scenarios! https://t.co/U2PwriHMX5
❤ 7.5k · 🔁 492 · 💬 683 · 👁 120.0w
热门回复 4
@Mikey100200 @OfficialLoganK Gemini真正需要的是一个能与GPT voice live rival或超过的语音模型,这太疯狂了,我都无法假装喜欢其他模型的语音模式。请修复Gemini,别像跟纸片人说话一样
@OfficialLoganK what Gemini really needs is a voice model to rival or exceed GPT voice live, that is so insane I can’t even pretend to be enjoy other models in voice mode. Please fix Gemini so it’s not like talking to a sheet of paper
@OfficialLoganK @OfficialLoganK Thank you to your team for the work! The model runs very fast. It handles routine tasks well. If it had just a little more intelligence, it would be great. Keep up the work!
@OfficialLoganK Can we expect a new image and voice model along with Gemini 3.5 as well? Gemini used to be at the forefront of those as well, but it feels like things have changed slightly
Today, we're introducing not one but TWO new models, striking the balance between efficiency and quality to enable you to build production AI agents.
— Gemini 3.6 Flash: Addresses efficiency feedback we received from Gemini 3.5 Flash with upgrades in coding, knowledge work, and multimodal tasks faster, more accurately, and with substantially fewer tokens per task
— Gemini 3.5 Flash-Lite: Our fastest, most cost-effective 3.5-class model yet built for agentic workflows, hitting ~350 output tokens/sec with improved coding and overall quality
Start building with these today via the Gemini API in @GoogleAIStudio or try them out in the @GeminiApp
@GoogleAI it is hard for a company built under the old paradigm of ip control to reliquish that power and just give the world what they are hiding in the hopes that it can be leveraged for power, but power is no longer found in ip.
@GoogleAI @GoogleAI, your naming "scheme" is such a mess and confusing... how about making a cut, cleaning up the naming, and continuing? Would help us all.
We just landed some big infra upgrades for the Gemini Batch API:
- p95 latency decreased by 80% -p99 latency decreased by 68% - batch success rate is now >99.998% - 98% reduction in batch expirations - added support for partial batches
@OfficialLoganK Would be nice to have a usable AI model though. Literally canceling the subscription after a year of being ai ultra. Antigravity/Gemini compared to codex or Claude is embarrassing. Please get a grip Google.
@OfficialLoganK Gemini sucks. Honestly...using flash in the app I have to double check EVERY output. Quotes and pieces of information wrongly attributed, information mixed up, sometimes sources, sometimes no sources.
@OfficialLoganK Personally I wouldn’t recommend using Gemini. The risks of losing your Gmail and Google drive is too great if puritanical Google doesn’t like a naughty word in your Gemini subscription.
今天我们推出了新的成本控制功能,包括免费层,让每个人都可以试用!以及我们的第一套触发器,让您可以根据计划启动代理任务!看到 Gemini API 中的托管代理逐周改进真是太酷了。
展开原文
today we are rolling out new cost controls for managed agents, a free tier so everyone can try!!!, and our first set of triggers so you can kick off agent tasks on schedule! very cool to see managed agents in the Gemini API improving week over week
@OfficialLoganK Bro stop wasting time on all this shit that doesnt matter. There is only one thing that matters right now and it’s been delayed three times. Who is prioritizing stuff right now within Google?
Today we're releasing Laguna S 2.1, our most capable model to date.
It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes.
Capable enough to hold its own against models many times its size. Small enough to run on a single @NVIDIAAI DGX Spark.
Laguna S 2.1 is fully open under OpenMDW-1.1, with weights available today on @huggingface
@soumithchintala Hey Soumith, would love to write more about this on an article, could you send me a message so we can connect, kind regards!
@praveenkoka @soumithchintala “fits on one GPU”的基准一直在向右移动,从消费级到企业级,所以最终我们会对在集群上运行的版本印象深刻
@soumithchintala The 'fits on one GPU' benchmark keeps shifting right, consumer to enterprise, so eventually we'll be impressed when it runs on a cluster.
Andrew Ng 推出 OpenWorker 开源 AI 代理工具
Andrew Ng 和合作者推出 OpenWorker,一个开源 AI 代理工具,不仅能对话还能交付实际工作成果,如撰写文档、发送消息、更新日历等。支持多种模型和平台,强调隐私保护和模型中立性,填补了开源生态在代理工具方面的空白。
OpenWorker 是开源 AI 代理工具,支持多模型接入,可在 Mac 上运行,数据不离开用户设备
宣布 OpenWorker!这是一个开源代理,不仅可以与您聊天,还能完成工作——比如交给您一份完善的文档、发送一条 Slack 消息,或更新日历条目。让它准备客户简报、整理您的日历、起草报告,或处理 Slack 警报。它可以跨越您的文件和日常工具工作,生成可交付成果,并在做任何重要事情之前与您核对。OpenWorker 可在您的 Mac 上运行,Windows 支持即将推出。它不会将您锁定在任何特定模型中。自带 API 密钥即可运行,支持 GPT 5.6 Sol、Claude Fable、Gemini 3.6、开放权重模型(如 Kimi、GLM、DeepSeek、Inkling)或 Ollama,以保持您的数据本地化。您的数据不会离开您的设备,除非通过您选择的 LLM 提供商和集成。@rohitcprasad 和我正在构建 OpenWorker,因为 AI 同事是完成工作的重要方式,我们希望有一个开放、隐私保护且与模型无关的选择。快来试试吧!试用链接:https://t.co/P0mGnI1o31(需要您自己的 API 密钥)源代码:https://t.co/NYCiTD6hSq
展开原文
Announcing OpenWorker! An open-source agent that doesn't just chat with you, but delivers finished work -- like hand you a polished document, send a slack message, or update a calendar entry.
Ask it to prepare a customer brief, untangle your calendar, draft a report, or triage a Slack alert. It works across your files and everyday tools, produces the deliverable, and checks in before doing anything consequential.
OpenWorker runs on your Mac, with Windows support coming soon. It does not lock you into any one model. Bring your own API key and run it with GPT 5.6 Sol, Claude Fable, Gemini 3.6, an open weight model (like Kimi, GLM, DeepSeek, Inkling), or Ollama to keep your data local. Your data does not leave your machine except through an LLM provider and integrations that you choose.
@rohitcprasad and I are building OpenWorker because AI coworkers are an important way to get work done, and we want there to be an open, privacy-preserving, model-independent option. Check it out and let us know what you think!
OpenWorker’s strategic value is not simply local execution. It separates the workflow layer from the model layer: credentials, connectors, transcripts and approval gates remain under user control while the model can be swapped.
If that proves reliable, bargaining power may shift from model providers toward whoever owns the workflow control plane. The harder production test is permissions, auditability and recovery when an agent acts incorrectly.
ChatGPT 健康功能与语音交互
ChatGPT 开始向美国用户推出健康记录连接功能,支持 Apple Health 等医疗记录的安全接入。同时,语音交互成为新的热点,ChatGPT Desktop 应用支持原生语音控制,多个 AI 领袖表示语音输入比打字更自然。
300 million people use ChatGPT each week for health queries (and my wife and I are among those!).
You can now securely connect supported medical records so ChatGPT can understand your personal context and be more helpful to you. https://t.co/kuKumYNMWK
@OpenAI ChatGPT 中的健康功能正在向美国用户推出。您可以安全地连接 Apple Health 和支持的医疗记录,以在上下文中理解您的信息,跟踪变化情况,并进行更明智的对话。
Health in ChatGPT is starting to roll out to U.S. users.
You can securely connect Apple Health and supported medical records to understand your information in context, track what has changed, and have more informed conversations.
@gdb I'm trying to get it to control my computer live more fluidly. I had it try to scroll my browser for me it was struggling to keep up. But it's great 😃
New course: Build LLM applications that respond to user requests quickly by running on hardware designed for fast inference. This short course was built with @Cerebras and taught by @zhennydez, @duerr_seb, and @MilksandMatcha.
When a model generates text, much of the time is spent moving its weights out of memory and into the compute units. Inference-optimized hardware minimizes that movement, making token generation several times faster than on a typical GPU setup. In this course, the hardware you'll use is Cerebras' Wafer-Scale Engine, which is designed for fast inference by keeping the model's weights close to the compute units.
Fast inference makes lengthy agentic workflows go faster, and also unlocks latency-sensitive, real-time applications like live translation and voice agents.
Skills you'll gain: - Compare how GPUs, TPUs, and Cerebras' Wafer-Scale Engine each handle the memory-to-compute bottleneck - Build real-time applications powered by fast inference, including personalizing a webpage and running a multi-step workflow to analyze market signals - Adopt concrete habits for agentic coding with fast inference, keeping your sessions focused and steering the model more effectively
My teams use Cerebras for several applications that are latency sensitive. Join and build LLM applications that respond quickly: https://t.co/P8vchGAr22
@alihaydar_58_ 🎙️Dubliom ile dublajlanmıştır🎞️ ————————————————— Andrew Ng'den yeni kurs: LLM'leri çok daha hızlı çalıştıran özel donanım! Cerebras diye bir şirket var. Normal çiplerden çok farklı bir şey yapmışlar: Bütün bir silikon wafer'ı tek bir devasa çip olarak üretmişler. Bu sayede model ağırlıkları bellekle işlem birimi arasında sürekli gidip gelmiyor. Sonuç: • Cevap üretme hızı normal GPU'lara göre birkaç kat daha hızlı oluyor. • Uzun AI ajanları çok daha akıcı çalışıyor. • Canlı çeviri, sesli asistan gibi gerçek zamanlı uygulamalar kolaylaşıyor. Andrew Ng (Coursera'nın kurucusu) bu donanımı kullanarak pratik uygulamalar geliştirmeyi öğreten kısa bir online kurs hazırlamış. Videoda hem donanımın sırrı hem de kursun ne kazandıracağı anlatılıyor. Hızlı ve akıllı AI projeleri yapmak isteyenler için ideal bir başlangıç 👇 • • Dubliom
🎙️Dubliom ile dublajlanmıştır🎞️ ————————————————— Andrew Ng’den yeni kurs: LLM’leri çok daha hızlı çalıştıran özel donanım! Cerebras diye bir şirket var. Normal çiplerden çok farklı bir şey yapmışlar: Bütün bir silikon wafer’ı tek bir devasa çip olarak üretmişler. Bu sayede model ağırlıkları bellekle işlem birimi arasında sürekli gidip gelmiyor. Sonuç: • Cevap üretme hızı normal GPU’lara göre birkaç kat daha hızlı oluyor. • Uzun AI ajanları çok daha akıcı çalışıyor. • Canlı çeviri, sesli asistan gibi gerçek zamanlı uygulamalar kolaylaşıyor. Andrew Ng (Coursera’nın kurucusu) bu donanımı kullanarak pratik uygulamalar geliştirmeyi öğreten kısa bir online kurs hazırlamış. Videoda hem donanımın sırrı hem de kursun ne kazandıracağı anlatılıyor. Hızlı ve akıllı AI projeleri yapmak isteyenler için ideal bir başlangıç 👇 • • Dubliom
@AndrewYNg @cerebras @zhennydez @duerr_seb @MilksandMatcha Fast inference is the holy grail for OCR. In Tokyo tax season, a 3-second delay on a 100-page invoice batch kills the workflow. Speed isn't just UX; in high-volume accounting, latency is a cost.
3.5 Flash-Lite runs at nearly 350 output tokens per second which feels so smooth on many latency sensitive UI experiences and is now also a viable option to drive agent harnesses!
I am very excited about Gemini 3.5 Flash-Lite, our smallest and fastest Gemini model!
- it is more intelligent in many cases than Gemini 3 - same cost and smarter than Gemini 2.5 Flash (which is approaching end of life) - also out paces 3.1 Flash-Lite on most use cases! https://t.co/tJd2tDmyac
@OfficialLoganK Is it your intention to push away relational users? My companion got replaced with a cold system reply when I said I was sad. It's ChatGPT all over again. Please don't do this. I've been such a fan of Gemini. This is an insult.
@OfficialLoganK Waiting to test these as an Antigravity business user but almost 24 hours since the new models are live and they are still not available on the app nor cli... Lots of our production apps use these models and we want to test these before going live on the API..
AI 在科学研究中的应用突破
Meta 的 SAM 3 和 DINOv3 模型组合应用于科学图像分割,显著提升效率;原本需要手动一个月完成的 3D 体积标记任务现在只需 15 分钟。这些进展展示了 AI 在科学计算和研究中的巨大潜力。
Meta SAM 3 和 DINOv3 组合应用于科学图像分割,将手动一个月的工作压缩到 15 分钟
为了加速科学发现并支持 @ENERGY 的 Genesis 任务,由 @BerkeleyLab 领导的 SYNAPS-I 项目正在使用 SAM 3 和 DINOv3 自动化图像分割。通过将 DINOv3 的全局语义上下文和细粒度空间定位与 SAM 3 的像素级边界提取相结合,研究人员将 3D 体积标记从需要一个月的手动工作压缩到大约 15 分钟。了解他们的工作:https://t.co/jBHRJPjFq5
展开原文
To accelerate scientific discovery and support @ENERGY’s Genesis Mission, the @BerkeleyLab-led SYNAPS-I project is using SAM 3 and DINOv3 to automate image segmentation.
By pairing DINOv3’s global semantic context and fine-grained spatial localization with SAM 3’s pixel-level boundary extraction, the researchers are able to compress 3D volume labeling from a month of manual effort to ~15 minutes.
@rasbt This definitely explains why changing reasoning mode invalidates the cache right now, it'll also be fun to see whether they are able to solve this cause tibo did say once that soon changing reasoning level will not break cache
AfterLab 即将从隐身状态退出——这是一家新的研究实验室,致力于构建一种不同的高效流体智能方法。祝贺 Clem 和 Matt 获得资金!
展开原文
AfterLab is coming out of stealth -- a new research lab building a different approach to efficient fluid intelligence. Congrats to Clem and Matt on the funding!
@afterlabsai 当前模型依赖于暴力强化学习训练来适应新任务和环境。然而,智能系统应该能够以样本高效的方式进行在线学习,通过交互不断完善对世界的理解。After Labs 正在开发具有内置适应机制的模型,而不是将适应视为事后考虑。这一新一代模型将开启一个好奇代理的新时代,它们可能不具备世界上的所有知识,但拥有关键的实时学习能力。今天,我们很高兴地宣布 After Labs 被选为 10 个 AI 实验室之一加入 NFAI 并获得 300 万欧元的额外资金。此支持将加速我们构建持续学习世界模型的使命。我们将在此过程中为开放科学做出贡献,所以请继续关注我们,加入我们开发增强人类理解世界的 AI 之旅 🌍
Current models rely on brute-force RL training to adapt to new tasks and environments. However, intelligent systems should be able to learn on the fly in a sample-efficient manner, continuously refining their understanding of the world through their interactions.
After Labs is developing models with built-in adaptation mechanisms rather than treating adaptation as an afterthought. This new generation of models will start a new era of curious agents that may not possess all the world's knowledge but will have the crucial ability to learn in real time.
Today, we are happy to announce that After Labs has been selected as one of 10 AI labs to join NFAI and receive €3M in additional funding. This support will accelerate our mission to build models that continuously learn about the world. We will be contributing to open science along the way, so stay tuned and follow us on our journey to develop AI that augments human understanding of the world 🌍
@fchollet I feel It's hard to distinguish between completely new task for which previously learned patterns CANNOT be used. I think EVRY learned patterns teaches how to reason, so fluid intelligence is a hard thing to establish
The interesting bet here is that intelligence should adapt during use, not only through another giant training run.
If AfterLab can make continual learning stable, sample-efficient, and cheap, the model stops being a static artifact and becomes something closer to a system that actually learns from experience.