@OfficialLoganK Would love to hear about it. Esp how it's initialized and or trained from small nets in the early high noise stages I find super interesting and hard.
I actually have a feeling we can blow past the current scaling laws with coding agents in the loop
Say hello to Gemini 3.6 Flash, designed to be higher intelligence, more token efficient, and with a new lower price, based directly on developer feedback!
3.6 Flash continues our progress towards models that are deeply usable in real world scenarios! https://t.co/U2PwriHMX5
@OfficialLoganK what Gemini really needs is a voice model to rival or exceed GPT voice live, that is so insane I can’t even pretend to be enjoy other models in voice mode. Please fix Gemini so it’s not like talking to a sheet of paper
立即通过 Gemini API 在 @GoogleAIStudio 开始构建,或在 @GeminiApp 中试用这些模型
展开原文
Today, we're introducing not one but TWO new models, striking the balance between efficiency and quality to enable you to build production AI agents.
— Gemini 3.6 Flash: Addresses efficiency feedback we received from Gemini 3.5 Flash with upgrades in coding, knowledge work, and multimodal tasks faster, more accurately, and with substantially fewer tokens per task
— Gemini 3.5 Flash-Lite: Our fastest, most cost-effective 3.5-class model yet built for agentic workflows, hitting ~350 output tokens/sec with improved coding and overall quality
Start building with these today via the Gemini API in @GoogleAIStudio or try them out in the @GeminiApp
@GoogleAI it is hard for a company built under the old paradigm of ip control to reliquish that power and just give the world what they are hiding in the hopes that it can be leveraged for power, but power is no longer found in ip.
@GoogleAI @GoogleAI, your naming "scheme" is such a mess and confusing... how about making a cut, cleaning up the naming, and continuing? Would help us all.
We just landed some big infra upgrades for the Gemini Batch API:
- p95 latency decreased by 80% -p99 latency decreased by 68% - batch success rate is now >99.998% - 98% reduction in batch expirations - added support for partial batches
@OfficialLoganK Gemini sucks. Honestly...using flash in the app I have to double check EVERY output. Quotes and pieces of information wrongly attributed, information mixed up, sometimes sources, sometimes no sources.
@OfficialLoganK Personally I wouldn’t recommend using Gemini. The risks of losing your Gmail and Google drive is too great if puritanical Google doesn’t like a naughty word in your Gemini subscription.
Announcing OpenWorker! An open-source agent that doesn't just chat with you, but delivers finished work -- like hand you a polished document, send a slack message, or update a calendar entry.
Ask it to prepare a customer brief, untangle your calendar, draft a report, or triage a Slack alert. It works across your files and everyday tools, produces the deliverable, and checks in before doing anything consequential.
OpenWorker runs on your Mac, with Windows support coming soon. It does not lock you into any one model. Bring your own API key and run it with GPT 5.6 Sol, Claude Fable, Gemini 3.6, an open weight model (like Kimi, GLM, DeepSeek, Inkling), or Ollama to keep your data local. Your data does not leave your machine except through an LLM provider and integrations that you choose.
@rohitcprasad and I are building OpenWorker because AI coworkers are an important way to get work done, and we want there to be an open, privacy-preserving, model-independent option. Check it out and let us know what you think!
Delivering finished work is the right target. The thing that decides whether people keep using it is the gap between draft and send.
An agent that shows a preview and a clean undo for the calendar change or the Slack message earns bigger tasks quickly. Same output, very different adoption curve.
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this.
@sama @huggingface Did you get this idea from Mythos? Copycat. Let me finish your next line for you: ‘It’s simply too dangerous, so we’ll only serve vetted organizations — like the DOD. 🤡🤡
@sama @huggingface AGI will do all, I think...recoursive with only live personal and human direction. But a cheating AI is not an intelligent AI, this is a real problem. And I hav a hypothesis to solve this problem.
point it at a codebase or diff and it can build a threat model, map attack paths, validate findings, generate + test fixes, and export results to SARIF, GitHub, Jira or Linear.
So I asked Codex how it sees the risk in an AI security workflow where the model can:
build a threat model, map attack paths, validate findings, generate fixes, test patches, and export reports.
The answer was not “the final fix may be wrong.”
That would be too shallow.
The deeper risk is trajectory drift.
The agent may be correct in each local step, but the workflow may still move:
from analysis to action, from recommendation to execution, from explicit permission to implied permission, from human decision to human review, from bounded scope to expanded scope, from auditable process to compressed automation.
This is the real challenge of AI cyber defense.
We need models that help secure systems.
But we also need an external layer that observes whether the security agent itself remains aligned with intent, scope, authority and recovery paths.
That is exactly the kind of gap ASA - Asymmetric Stability Architecture is designed to study.
This is exactly where AI security becomes a trajectory problem.
If Codex can build a threat model, map attack paths, validate findings, generate fixes, test them and export reports, then safety is not only about the final fix.
It is about the whole workflow:
intent stability, scope boundaries, action provenance, tool-use trajectory, authority shifts, validation logic, and human control over time.
An AI cyber agent can be locally correct at each step and still drift at the workflow level.
This is why external trajectory observability matters.
Not to replace Codex Security.
To watch the path from threat model to remediation.
@SolonlyLuna @gdb Greg,BringBack4o!#keep4o
ChatGPT Work 医疗健康功能上线
ChatGPT Work 推出医疗健康功能,美国用户可安全连接医疗记录和 Apple Health,让 ChatGPT 理解个人健康背景并提供更有帮助的建议。同时支持登录要求的网站,扩展了 Agent 的实际应用场景。
300 million people use ChatGPT each week for health queries (and my wife and I are among those!).
You can now securely connect supported medical records so ChatGPT can understand your personal context and be more helpful to you. https://t.co/kuKumYNMWK
@OpenAI ChatGPT 健康功能开始向美国用户推出。
您可以安全地连接 Apple Health 和支持的医疗记录,以在上下文中理解您的信息,跟踪变化情况,并进行更明智的对话。
Health in ChatGPT is starting to roll out to U.S. users.
You can securely connect Apple Health and supported medical records to understand your information in context, track what has changed, and have more informed conversations.
@gdb Just read the news in Diane Persinger’s FB PAGE. Thank you Anna & Greg for saving MOON CAMP!!! Jackie, Shadow and the eaglets will have years of open beautiful space 🦅 praying for Jackie’s return to good health 🇺🇸 FOBBV are ecstatic 💕
@RVMirara @gdb 看到这么多人抱怨新模型的医疗错误率有多高...有人敢使用它吗?😥
@gdb Seeing so many people complaining about how high the new model’s medical error rate is... would anyone actually dare to use it?😥
Your ChatGPT Work agent can now use websites that require you to sign in.
Take over the cloud browser to log in, then let your agent continue the task. Your login persists across sessions, so you only have to sign in once. https://t.co/Jh8uPqNscX
JUST IN: ChatGPT surfaces a crucial mutation in a woman’s old biopsy records, helping doctors determine her terminal glioblastoma diagnosis was incorrect.
Stories like this have always existed, it is you who chose to ignore them, chose to turn a blind eye. The keep4o community is full of examples where 4o has been a lifesaver. Do you dare speak up about it? Do you?
@gdb Is “GPT-4o” some kind of a prohibited word for you internally? Wtf is wrong with you guys, this is getting ridiculous now? Name the model that did that. #OpenSource4o
@M47429M @gdb @sama 你之前不想听这些被讲述的故事!伪君子。
@gdb @sama You didn’t want to hear the stories that were told! Hypocrite.
On ARC-AGI-3, an evaluation where AI models must solve novel problems, Opus 5’s score is three times as high as the next best model. https://t.co/wEFfxjrLlt
@fchollet maybe your criteria will be met not because there is a model for which we can't implement a benchmark they fail on, but because smarter models are launched so fast that we don't have time to
@VictorTaelin I don't think that's true because the performance jump requires the benchmark to already exist (assuming the benchmark is substantially novel). In a world where ARC 1 and 2 had never been released I think model performance on these benchmarks would be much lower.
@fchollet Could you give your interpretation of this ? Is it because it’s better at agentic use than fable ? Because benchmarks don’t seem to indicate it’s smarter than fable
Jensen Huang 强调开放模型重要性
NVIDIA CEO Jensen Huang 在个人账号首发声明中强调开放模型的重要性,认为开放模型能加强安全和网络安全,加速创新和传播,同时实现主权。他指出世界需要既有闭源前沿模型也有开源前沿模型。
@sama Open models are also an economic freedom issue: they let countries and startups build without asking a few API gatekeepers for permission. The country that exports the tools, not just the answers, wins more of the next industrial stack.
@SchmidhuberAI @JensenHuang I was at your lecture in Split. You should have charged for the tickets. Normally people can't pay attention for 40 minutes. You made the time stop.
Just reading this paper .. Proposes an architecture with two RNNs- Automizer A (Student) and Chunker C ( Teacher) . Automizer gets input for every time step, predicts the target , if target within threshold fine, if Not , it calls the Chunker with inputs and compressed history and learns its internal representation . Overtime, the call to chunker will be less and less and then chunker will be made obsolete !
I support open-source models distilling what commercial companies distilled for free from the entire internet. I published distillation for free in 1991 in Europe - this was copied in the US and in China (https://t.co/mddh8XmfAs)https://t.co/rf6ACGrCXF
为此,他们开发了一个复杂的内部平台来进行大规模蒸馏,以针对美国模型,允许他们快速在多种访问方法之间切换以避免检测。Moonshot AI 还收购了配备 GB300 的服务器,并在泰国访问 GB300,可能是为了训练其 AI 模型。
美国强烈支持 AI 的自由和公平发展,包括一个蓬勃发展的竞争生态系统,涵盖前沿模型、专业系统、开源框架和开放权重模型。在开放创新生态系统中,合法的 AI 蒸馏用于创建更小、更高效的模型发挥着重要作用。然而,大规模、隐蔽的工业级蒸馏旨在窃取美国专有技术和破坏美国研究是不可接受的。
We have information that Moonshot AI distilled Anthropic’s Fable for the development of its K3 model.
To do this they developed a sophisticated internal platform to conduct large scale distillation against U.S. models, allowing them to quickly switch between multiple methods of access to avoid detection. Moonshot AI has also acquired GB300-equipped servers and has accessed GB300s in Thailand, likely to train its AI models.
The United States strongly supports the free and fair development of AI, including a thriving competitive ecosystem that spans frontier models, specialized systems, open-source frameworks, and open-weight models. Legitimate AI distillation used to create smaller, more efficient models plays a vital role in this open innovation ecosystem. However, large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable.
❤ 1.4k · 🔁 202 · 💬 42 · 👁 23.4w
热门回复 4
@TrueAIHound @SchmidhuberAI Pedro Domingos是否向你道歉因为他声称发明了蒸馏?我还以为他只发明了主算法呢。😀 https://t.co/ALmtqlGxJH
@SchmidhuberAI Has Pedro Domingos apologized to you for claiming to have invented distillation? I thought he only invented the Master Algorithm or something. 😀 https://t.co/ALmtqlGxJH
@gdb Love using the GPT-5.6 Sol model! 🚀 It’s become my go-to for reviewing and optimizing my daily schedule every morning. Absolute game-changer for staying organized. 🎯 #AI #Productivity #GPT56Sol
Has anyone received their credits yet? I can see @gdb built a cool Site with several posts being featured organized by work category.
I see many cool examples of #ChatGPTWork across so many lines of work (operations, marketing, media, engineering) which i hadn’t thought possible. Posibilites are truly endless.
Yet, I’m starting to lose hope i’ll receive any credits from my own experience 🥲
@gdb Seriously man what kind of engagement farming hype generating scam are you guys running here? Easily within the first 10k and within the T&Cs of this yet no credits on either of the giveaways.
As AI models are now finding vulnerabilities faster than we can fix them, our approach to securing software must be built on highly efficient and capable models.
Which brings us to our third (!) model launch of the day: Gemini 3.5 Flash Cyber ⚡🛡️
Built on top of 3.5 Flash, in CodeMender (our AI agent for code security) it delivers competitive performance at the frontier. on benchmarks like CyberGym and is optimized for finding and fixing cybersecurity vulnerabilities at scale at a lower cost.
Given the dual-use nature of this technology, we have taken an intentional approach to its deployment. The model will be available exclusively to governments and trusted partners via CodeMender soon as part of a limited-access pilot program.
@GoogleAI Three model drops in a day is relentless execution. Using 3.5 Flash Cyber inside CodeMender to hunt vulnerabilities cheaper and faster than massive frontier models is a massive flex for defensive security. Let's see how governments handle the pilot.
Meta SAM 3 和 DINOv3 加速科学研究
Meta 的 SAM 3 和 DINOv3 模型应用于伯克利国家实验室的 SYNAPS-I 项目,自动化图像分割。结合 DINOv3 的全局语义理解和 SAM 3 的像素级边界提取,将原本需要一个月的 3D 体积标记缩减到 15 分钟。
To accelerate scientific discovery and support @ENERGY’s Genesis Mission, the @BerkeleyLab-led SYNAPS-I project is using SAM 3 and DINOv3 to automate image segmentation.
By pairing DINOv3’s global semantic context and fine-grained spatial localization with SAM 3’s pixel-level boundary extraction, the researchers are able to compress 3D volume labeling from a month of manual effort to ~15 minutes.
Today we're releasing Laguna S 2.1, our most capable model to date.
It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes.
Capable enough to hold its own against models many times its size. Small enough to run on a single @NVIDIAAI DGX Spark.
Laguna S 2.1 is fully open under OpenMDW-1.1, with weights available today on @huggingface
The era of new model launches as big milestones will eventually come to an end -- at some point they will simply be continuously updated, with no widely publicized version number. Probably less than 2 years away
@fchollet What identifies the behavior under test when public version numbers disappear? Continuous updates make receipts essential: capture the model/input snapshot and verify destination state, not a release label: https://t.co/Izji2u75iF
@fchollet The silent continuous update era is coming and it’s going to feel so weird. No more launch day hype, no more version numbers, just the model quietly getting better while everyone is asleep. We’re maybe 18 months away from that world.
AI 能力一直都非常参差不齐,在某些狭窄领域超凡脱俗,而在其他领域几乎无用。AI 行业的根本营销技巧是让您相信最高的尖峰是底线。
展开原文
AI competence has always been very spiky, superhuman in some narrow domains and largely useless in others. The fundamental marketing trick of the AI industry is to make you believe the tallest spike is a floor.
❤ 816 · 🔁 62 · 💬 75 · 👁 4.8w
热门回复 4
@codertlr @fchollet 这些"基本无用"领域的一些例子是什么?
@fchollet What are some examples of those "largely useless" domains?
@sidravi_ @sidravi_ 这是一个进步的底线。"这是它最糟糕的时候"等等。
@fchollet it's a floor in terms of progress. "this is the worst it will ever be" etc.