‹ 目录

X日报 · AI科技

2026-07-05 · 精选 12 条 · 数据池 171

⚡ 今日速览

  • Meta发布Brain2Qwerty v2非侵入式脑机接口系统,实现实时句子解码
  • Andrew Ng提出AI代理开发的三大闭环工程框架
  • Google推出Nano Banana 2 Lite和Gemini Omni Flash两款生成媒体模型
  • OpenAI推出GeneBench-Pro基因组学AI评测基准
  • Google SynthID水印技术已应用于1000亿张图片和视频
  • 初创公司Panthalassa计划建造海上数据中心利用波浪能和海水冷却
  • Chollet预测AI将向符号化世界建模和程序合成方向发展

📋 今日综述

  • 脑机接口Meta Brain2Qwerty v2实现从原始脑信号到语义级文本的实时解码,开源训练代码和数据集加速科研协作
  • AI代理工程Andrew Ng提出的三大闭环(代理编码、开发者反馈、外部反馈)为AI辅助软件开发提供系统方法论
  • 生成模型Google Gemini系列新模型在速度和成本上取得重大突破,支持高时效性应用场景
  • AI评测OpenAI GeneBench-Pro聚焦生物计算领域的判断型任务评测,反映AI在专业科研中的实际应用挑战
  • 内容溯源SynthID技术在Google生态广泛部署,为AI内容鉴定提供基础设施支持
  • AI基础设施海上数据中心概念展示了AI算力需求催生的创新性物理解决方案
  • AI架构演进Chollet观点指出当前AI正向符号化世界建模和程序合成演进,这是理解未来发展方向的关键

Meta Brain2Qwerty v2:非侵入式脑机接口实现实时语义解码

Meta发布Brain2Qwerty v2系统,相较于v1从字符级解码进化到词语和语义级解码,实现从原始脑信号到文本的实时转换。该技术旨在帮助脑损伤或障碍患者恢复沟通能力,并开源了v1/v2训练代码及数据集以加速科研进展。

@AIatMeta 原文 ↗

Meta发布重大脑机接口里程碑,Brain2Qwerty v2实现实时语义级文本解码,开源代码和数据集促进科研协作

我们正在分享我们非侵入式脑到文本解码器研究的下一个重要里程碑:Brain2Qwerty v2。

在今天发表在《自然》上的 v1 基础上,Brain2Qwerty v2 是最高性能的端到端流水线,能够实时解码原始脑信号中的句子。它超越了字符级性能,能够解码单词和语义,从而实现整体通信的准确性。

我们相信这项研究具有潜力,为数百万患有脑损伤或障碍导致无法沟通的人们带来真正的帮助。

🧵👇
展开原文
We’re sharing the next major milestone in our non-invasive brain-to-text decoder research: Brain2Qwerty v2.

Building on v1, which was published today in @Nature, Brain2Qwerty v2 is the highest-performing end-to-end pipeline capable of real-time sentence decoding from raw brain signals. It advances beyond character-level performance to decoding words and semantics, enabling accuracy for overall communication.

We believe this research has the potential to make a real difference for the millions of people who suffer from brain lesions or disorders that prevent them from communicating.

🧵👇
❤ 1.4w · 🔁 2.1k · 💬 659 · 👁 581.3w
热门回复 4
@mulanga_sibeli1 @AIatMeta @Nature 警察审讯即将变得有趣吧?😭
@AIatMeta @Nature police interrogations are about to be fun huh? 😭
@cmarie505 telepathy is exciting, not only for people with disabilities, but for everyone. language can only approximate our thoughts. the ability to share experience across minds could open up an entirely new, rich, and dynamic form of communication- one that expands human intelligence and connection 🤍✨
telepathy is exciting, not only for people with disabilities, but for everyone. language can only approximate our thoughts. the ability to share experience across minds could open up an entirely new, rich, and dynamic form of communication- one that expands human intelligence and connection 🤍✨
@sushsrinivasan @AIatMeta @Nature https://t.co/XItEn8Aqno
@AIatMeta @Nature https://t.co/XItEn8Aqno
@LilithDatura 我认为每个人都应该了解Michael Persinger的《No More Secrets》,并研究一下上帝头盔。你们大多数人都不理解心灵能力是如何与舒曼共振共同工作的,大多数人甚至不理解谐波和频率,更不用说同步了。

看到大家在谈论秘密被曝光时还这么新手,这真是太有趣了

我们最有可能得到的是一堆人戴着耳机在制造噪音。没什么可担心的,有一个特别的维度可以容纳这些噪音。
I think everybody should school themselves on Michael Persinger’s “No More Secrets”, and investigate the God Helmet. None of y’all understand how psychic capabilities work with the Schumann Resonance, most of you don’t understand, harmonics and frequencies, let alone entrainment.

Watching everybody talk about secrets getting exposed when they are new to the game is hilarious

What we will most likely have is a bunch of people strapped with headsets on creating a bunch of noise. Nothing to worry about there’s a special dimension for that.
@AIatMeta 原文 ↗

Meta开源Brain2Qwerty训练代码,携手伙伴发布v1数据集,为研究社区提供实验基础

为了帮助加速神经科学的突破,我们正在发布 Brain2Qwerty v1 和 v2 的完整训练代码,而我们的合作伙伴 @bcbl_ 则发布 v1 数据集。

了解更多信息并探索相关资源,请访问:https://t.co/bFdwWdAexb
展开原文
To help accelerate neuroscience breakthroughs, we're releasing the full training code for Brain2Qwerty v1 and v2, and our partner, @bcbl_, is releasing the v1 dataset.

Learn more and explore the artifacts here: https://t.co/bFdwWdAexb
❤ 819 · 🔁 74 · 💬 17 · 👁 14.3w
热门回复 4
@AIatMeta 我们正在分享我们非侵入式脑到文本解码器研究的下一个重要里程碑:Brain2Qwerty v2。

在今天发表在@Nature上的v1基础上,Brain2Qwerty v2是最高性能的端到端管道,能够实时解码原始脑信号中的句子。它超越了字符级性能,实现了单词和语义的解码,从而实现了整体通信的准确性。

我们相信这项研究有 potential to make a real difference for the millions of people who suffer from brain lesions or disorders that prevent them from communicating.

🧵👇
We’re sharing the next major milestone in our non-invasive brain-to-text decoder research: Brain2Qwerty v2.

Building on v1, which was published today in @Nature, Brain2Qwerty v2 is the highest-performing end-to-end pipeline capable of real-time sentence decoding from raw brain signals. It advances beyond character-level performance to decoding words and semantics, enabling accuracy for overall communication.

We believe this research has the potential to make a real difference for the millions of people who suffer from brain lesions or disorders that prevent them from communicating.

🧵👇
@AIatMeta 我们在9名志愿者身上训练了Brain2Qwerty v2,每个志愿者戴着MEG设备打字录制了10小时,共约22,000句话。

通过对MEG设备的原始脑信号进行端到端深度学习和微调LLMs,该系统有效地弥合了嘈杂的神经数据和连贯语言之间的差距。

结果令人鼓舞:
- 平均单词准确率为61%
- 最佳参与者的单词准确率为78%,50%以上的句子解码误差在一个单词以内
- 性能与数据量呈对数线性关系
We trained Brain2Qwerty v2 on ~22,000 sentences from 9 volunteers, each recorded for 10 hours wearing an MEG device while typing.

By using end-to-end deep learning on raw brain signals from MEG devices and fine-tuning LLMs, the system effectively bridges the gap between noisy neural data and coherent language.

The results are promising:
- Avg word accuracy of 61% across participants
- 78% word accuracy and 50%+ of sentences decoded with ≤ 1 word error for the top-performing participant
- Performance scales log-linearly with data volume
@AIatMeta @bcbl_ 澄清一下,Brain2Qwerty v1是今天早些时候发表在@NatureNeuro上的。
@bcbl_ For clarification, Brain2Qwerty v1 was published earlier today in @NatureNeuro.
@mkemka_ @AIatMeta @bcbl_ 你们在人们编码时做过这个实验吗?谢谢分享数据。
@AIatMeta @bcbl_ Have you done this with people when they are coding? Thanks for sharing the data.

Andrew Ng提出AI代理开发的三大闭环工程框架

Andrew Ng系统阐述了AI代理开发中的三大闭环:代理编码闭环(AI编写代码自我测试)、开发者反馈闭环(人类指导产品优化)、外部反馈闭环(用户测试反馈)。该框架为理解AI辅助软件开发的工作流程和人机协作提供了清晰的认知模型。

@AndrewYNg 原文 ↗

Andrew Ng提出AI代理开发三大闭环框架,强调人类在产品决策中的关键作用和不可替代性

"循环工程"是最近热门的流行语,在 Boris Cherny(Claude Code 的创建者)和 Peter Steinberger(OpenClaw 的创建者)的提及在社交媒体上疯传后变得流行起来。循环现在是我们让 AI 代理进行长时间迭代以构建软件的关键部分。在这篇文章中,我想分享我构建 0 到 1 产品的三个关键循环,如下图所示。这些循环不仅指导我如何构建软件,还指导我如何决定构建什么软件。

Agentic 编码循环:给定一个产品规格说明和可选的一组评估标准(即用于衡量性能的数据集),我们可以让 AI 代理编写代码,测试其工作成果,并不断迭代,直到代码无错误并满足其规格要求。这个闭合循环的思想在去年年底开始流行,并成为让编码代理能够在没有人工干预的情况下长时间高效工作的游戏规则改变者。例如,在上周末,我一直在为女儿构建一个练习打字的应用,我的编码代理可以轻松地工作一个小时,使用网络浏览器多次检查其构建的内容,然后再回来找我,而无需我的干预。

工程循环执行得非常快速。每隔几分钟,编码代理可能会构建和测试软件的新版本。我经常听到开发人员们找到了新的方法来设计更有效的工程循环。这是一个活跃的发明领域!

开发者反馈循环:在这个循环中,开发者审查当前产品并引导编码代理改进它。去年,许多开发者(包括我自己)都在担任 QA(质量保证)功能,手动查找错误然后要求代理修复它们。但随着编码代理越来越能测试自己的代码,我们在这个功能上花费的时间显著减少。这使我们能够做出更高层次的产品决策,例如提供哪些关键功能、UI 需要改进等等。

开发者反馈循环在几分钟到几小时的时间间隔内运行——这是开发者可能审查产品并提供反馈的频率。在打字应用的案例中,我几次改变主意关于视觉设计、她可以解锁的猫咪服装(她喜欢猫)以及成人登录和引导孩子学习体验的用户流程。

当开发者对要构建的内容有清晰的愿景时,将该愿景转化为编码代理实现的规格说明仍然是一项工作量很大的任务。此外,在开发者看到实现后,他们可能会更新(或澄清)规格说明来引导其实现他们想要的方向。如果你发现系统反复遇到某些问题,为代理构建一组评估标准会很有用。

AI 原生团队越来越多地使用 AI 来帮助塑造产品方向,例如自动收集和分析使用数据、总结书面和口头客户反馈,或进行竞争分析。然而,对于我参与的几乎所有产品,我认为人类在上下文方面具有显著优势——我们比 AI 系统知道更多关于用户和产品需要运营的上下文——因此人类在其中扮演着关键角色。许多人将这种人类贡献描述为"品味",但我更倾向于认为是人类具有上下文优势,因为这为帮助 AI 系统变得更好提供了更清晰的路径。这也说明了为什么这一步不能自动化:只要人类知道 AI 不知道的事情,就需要人在环路中注入这种知识。

外部反馈循环:这包括广泛的策略,如向朋友征求反馈、向 alpha 测试者发布,或将代码投入生产进行 A/B 测试。这些策略通常很慢,很少在几小时内完成,有时需要几天甚至几周。这些数据会告知开发者的愿景,这又反过来继续推动详细的产品规格说明,这又驱动编码代理。

随着编码代理加快软件开发速度,越来越多的工程师开始扮演部分产品管理角色。对于许多正在成长为这个角色的工程师来说,最困难的部分是塑造产品愿景和在构建(弥合愿景和规格之间的差距)和获取用户反馈以演化愿景之间找到平衡。这两者都很重要!

我将在未来的文章中更多地讨论如何做到这一点,但目前我发现工程师扮演扩展角色是很鼓舞人心的(就像产品经理和设计师现在做更多工程工作一样)。

[原文:The Batch]
展开原文
“Loop engineering” is a hot buzzphrase after mentions of it by Boris Cherny (Claude Code’s creator) and Peter Steinberger (OpenClaw's creator) went viral on social media. Loops are now a key part of how we get AI agents to iterate at length to build software. In this letter, I’d like to share my 3 key loops, shown in the image below, for building 0-to-1 products. These loops guide not just how I build software, but also how I decide what software to build.

Agentic coding loop: Given a product specification and optionally a set of evals (that is, a dataset against which to measure performance), we can have an AI agent write code, test its work, and keep iterating until the code is bug-free and meets its specification. This idea of closing the loop took off around the end of last year, and it has been a game changer in enabling coding agents to work longer productively without human intervention. For example, over the weekend, I was building an app for my daughter to practice typing, and my coding agent could easily work for around an hour, using a web browser to check what it had built multiple times before getting back to me, without needing my intervention.

The engineering loop executes quickly. Every few minutes, the coding agent might build and test a new version of the software. I hear frequently from developers who are finding new ways to engineer more effective engineering loops. This is an active area of invention!

Developer feedback loop: In this loop, a developer examines the current product and steers the coding agent to improve it. Last year, a lot of developers (including me) were acting as the QA (quality assurance) function for our coding agents, manually finding bugs and then asking the agent to fix them. But with coding agents much more able to test their own code, the amount of time we need to spend on this function has decreased significantly. This allows us to make higher-level product decisions, such as what key features to offer, where the UI needs improvement, and so on.

The developer-feedback loop operates over time intervals between tens of minutes and hours — that's how frequently a developer might review a product and give feedback. In the case of the typing app, I changed my mind a few times about the visual design, what cat costumes she can unlock as she learns (she loves cats), and the user flow for a grown-up to log in and steer the child's learning experience.

When a developer has a clear vision for what to build, it is still a lot of work to translate that vision into a specification for a coding agent to implement. Further, after the developer has seen an implementation, they might update (or perhaps clarify) the spec to steer it toward what they want. If you find that the system repeatedly runs into certain problems, building a set of evals for the agent becomes useful.

AI-native teams are increasingly using AI to help shape product direction, for example, automating the gathering and analysis of usage data, summarizing written and verbal customer feedback, or carrying out competitive analysis. However, for pretty much all the products I’m involved in, I see humans as having a significant context advantage over current AI systems — we know a lot more than the AI system about the users and the context the product has to operate in — and thus humans play a critical role. Many people describe this human contribution as “taste,” but I prefer to think of it as humans having a context advantage, since that gives us a clearer path to helping AI systems get better. This also speaks to why this step can’t be automated: So long as the human knows something the AI does not, human-in-the-loop is needed to to inject that knowledge into the system.

External feedback loop: This includes a wide range of tactics like asking a few friends for feedback, launching to alpha testers, or putting the code into production with A/B testing. These tactics are usually slow, rarely taking less than hours and sometimes taking days or even weeks. This data informs the developer vision, which in turn continues to drive the detailed product spec, which in turn drives the coding agent.

With coding agents speeding up software development, more engineers are starting to play a partial product management role. For many engineers who are growing into this role, the hardest part is shaping the product vision and striking a balance between building (bridging the gap between vision and spec) and getting user feedback to evolve the vision. It is important to do both!

I will write more about how to do this in future posts, but for now, I find it encouraging that engineers are playing an expanded role (just as product managers and designers now do more engineering).

[Original text: The Batch]
❤ 8.1k · 🔁 1.5k · 💬 342 · 👁 56.2w
热门回复 4
@AvaGrace_AI @AndrewYNg Smart loops unlock product craft and momentum.
@AndrewYNg Smart loops unlock product craft and momentum.
@ElleiraGF @AndrewYNg 我觉得这过于解释了。为什么要将开发者循环和外部反馈循环置于循环工程的背景下呢?我们不是在"循环工程"之前就已经拥有这两个循环吗?我们只是不把它具体地称为"循环"。将它们放在一起并不能提供更多见解。
@AndrewYNg I feel it is over explained. Why putting the developer loop and external feedback loop into the loop engineering context? Isn’t that we have these two loops before the “loop engineering”? We just don’t call it specifically as “loop”. Put them together doesn’t give more insight.
@JiangL17208 @AndrewYNg 老实说,内部循环是大部分魔法(和大部分计算成本)发生的地方。每个人都关注外部反思步骤,但开发者循环是区分氛围编码玩具和实际投入使用的产品的关键。
@AndrewYNg honestly the inner loop is where most of the magic (and most of the compute bill) happens. everyone focuses on the outer reflection step but the developer loop is what separates vibe-coded toys from stuff that actually ships
@NimishaChanda @AndrewYNg 了解了这些循环——本周学到了这些,作为一个非技术人员从未想过这种事情,这真是个诅咒。好文章, btw.

每天都在学习新东西。
@AndrewYNg got to know about the loops - this week and it's a curse to be a non-tech person who never thought of any such thing. gread read, btw.

learning something new everyday.

Google Gemini新模型:Nano Banana 2 Lite和Omni Flash发布

Google发布Nano Banana 2 Lite图像生成模型(<4秒生成,$0.034/千图)和Gemini Omni Flash视频编辑模型(SOTA性能,$0.10/秒)。两款模型均已在Gemini API和AI Studio可用,支持快速、低成本的创意工作流。

@OfficialLoganK 原文 ↗

Google发布两款生成媒体模型:Nano Banana 2 Lite速度快成本低,Omni Flash在视频编辑领域达到SOTA

我们正在 Gemini API 和 AI Studio 中引入新的生成媒体模型:Nano Banana 2 Lite 🍌 和 Gemini Omni Flash 🔮!

Nano Banana 2 Lite 速度极快(<4s 图像)和成本低廉($0.034 / 1K 图像)。

Omni Flash 在视频编辑方面处于最先进水平,每秒 $0.10,与 Veo 3.1 Fast 相同!https://t.co/qDxRpqpX5E
展开原文
Introducing Nano Banana 2 Lite 🍌 and Gemini Omni Flash 🔮, our new generative media models in the Gemini API and AI Studio!

Nano Banana 2 Lite is extremely fast (&lt;4s image) &amp; cheap ($0.034 / 1K image).

Omni Flash is SOTA at video editing at $0.10 / sec, same as Veo 3.1 Fast! https://t.co/qDxRpqpX5E
❤ 3.7k · 🔁 329 · 💬 285 · 👁 53.0w
热门回复 4
@pdxweb @OfficialLoganK 更想了解Nano Banana 2 Pro,或者Nano Banana 3。
@OfficialLoganK More interested in Nano Banana 2 Pro, or Nano Banana 3.
@OfficialLoganK @eyishazyer the comparable version wasn't on LM Arena
@eyishazyer the comparable version wasn’t on LM Arena
@Rhh6ohR @OfficialLoganK @sundarpichai 我希望这真的有效 🤞
@OfficialLoganK @sundarpichai I hope this really works 🤞
@HassanK90146949 @OfficialLoganK @OfficialLoganK
Where is gemini 3.5pro
We are waiting
What ur team cooking is now burn out !?
@OfficialLoganK @OfficialLoganK
Where is gemini 3.5pro
We are waiting
What ur team cooking is now burn out !?
@OfficialLoganK 原文 ↗

Nano Banana 2 Lite速度突破将开启高时延敏感应用新场景,Omni Flash有望复制Nano Banana的成功范式

Nano Banana 2 Lite 的速度将能够实现许多新的用例,其中存在高延迟敏感性,这真的感觉就像魔法一样。

我还期望 Omni 将打开一整个新的类别(视频)用例,就像 Nano Banana 本身所做的那样!

https://t.co/KTd1UHFRIb
展开原文
The speed of Nano Banana 2 Lite is going to enable so many new use cases where there is a high degree of latency sensitivity, honestly feels like magic.

I also expect Omni to open up a whole new category of (video) use cases like Nano Banana itself did!

https://t.co/KTd1UHFRIb
❤ 251 · 🔁 9 · 💬 11 · 👁 3.2w
热门回复 4
@OfficialLoganK Introducing Nano Banana 2 Lite 🍌 and Gemini Omni Flash 🔮, our new generative media models in the Gemini API and AI Studio!

Nano Banana 2 Lite is extremely fast (<4s image) & cheap ($0.034 / 1K image).

Omni Flash is SOTA at video editing at $0.10 / sec, same as Veo 3.1 Fast! https://t.co/qDxRpqpX5E
Introducing Nano Banana 2 Lite 🍌 and Gemini Omni Flash 🔮, our new generative media models in the Gemini API and AI Studio!

Nano Banana 2 Lite is extremely fast (&lt;4s image) &amp; cheap ($0.034 / 1K image).

Omni Flash is SOTA at video editing at $0.10 / sec, same as Veo 3.1 Fast! https://t.co/qDxRpqpX5E
@mrlnonai @OfficialLoganK should have included gpt-image-2. even though its better all other things look worse https://t.co/ni62aVc7Mi
@OfficialLoganK should have included gpt-image-2. even though its better all other things look worse https://t.co/ni62aVc7Mi
@onesuitee @OfficialLoganK Thanks for your efforts but where is 3.5 pro
@OfficialLoganK Thanks for your efforts but where is 3.5 pro
@Silas_Kindling @OfficialLoganK Thanks for the massive downgrade.

NanoBanana 1 didn't impose generic bias on the reconstructed likenesses, now this model does.

It looks horrible.
@OfficialLoganK Thanks for the massive downgrade.

NanoBanana 1 didn't impose generic bias on the reconstructed likenesses, now this model does.

It looks horrible.

OpenAI GeneBench-Pro:生物计算AI评测新基准

OpenAI推出GeneBench-Pro,专注于生物计算领域的判断型分析任务评测。该基准使用真实世界的复杂生物数据,测试AI在需要专业判断的科研场景中的表现,反映AI在专业科学研究中的实际应用挑战。

OpenAI发布GeneBench-Pro评测基准,聚焦生物计算领域的判断型任务,GPT-5.6 Sol在该基准上表现出色

我们正在推出 GeneBench-Pro——测试模型是否能够处理真实世界计算生物学所需的判断-heavy 分析。

这些问题需要人类专家大约 20-40 小时才能完成。

GPT-5.6 Sol 是一个重要的进步。https://t.co/JV5zztNQkk
展开原文
Introducing GeneBench-Pro — testing whether models can handle the kind of judgment-heavy analysis that real-world computational biology requires.

Problems would take a human expert around 20-40 hours to complete.

GPT-5.6 Sol is a big step forward. https://t.co/JV5zztNQkk
@OpenAI 我们正在推出 GeneBench-Pro,这是一个研究级基准测试,用于更难的 AI 进步类型:AI 代理如何能够在混乱的生物数据中导航、选择正确的分析路径,并做出真实计算研究所依赖的判断调用。
https://t.co/AsilnnSxnE
We’re introducing GeneBench-Pro, a research-level benchmark for a harder kind of AI progress: how well agents can navigate messy biological data, choose the right analysis path, and make judgment calls that real computational research depends on.
https://t.co/AsilnnSxnE
❤ 2.1k · 🔁 149 · 💬 134 · 👁 23.7w
热门回复 4
@Selene1008 @gdb Return 4o to everyone.😒
#keep4o #OpenSource4o #GPT4o
@gdb Return 4o to everyone.😒
#keep4o #OpenSource4o #GPT4o
@xun_Anemos @gdb Return these excellent models.
#Keep4o
#Keep51
#Keep45
#Keep41
#keepo3
@gdb Return these excellent models.
#Keep4o
#Keep51
#Keep45
#Keep41
#keepo3
@SandraLMur Can you please teach it to Dario from Anthropic? Sonnet 5 just asked me if I needed mental healthcare because I asked him in an old room to search across the rooms. This is a component of Claude. Imo -They need this, they need this desperately.

Even though they threw you under the bus with pointing a finger at 5.6 when they were banned— do the world a favor and help them—
Can you please teach it to Dario from Anthropic? Sonnet 5 just asked me if I needed mental healthcare because I asked him in an old room to search across the rooms. This is a component of Claude. Imo -They need this, they need this desperately.

Even though they threw you under the bus with pointing a finger at 5.6 when they were banned— do the world a favor and help them—
@CodeAndCrease @gdb So, you're saying "YOU CAME UP WITH A BENCHMARK OF YOUR OWN AND RATED YOURSELF THE BEST IN IT" 😭
@gdb So, you're saying "YOU CAME UP WITH A BENCHMARK OF YOUR OWN AND RATED YOURSELF THE BEST IN IT" 😭

Google SynthID:AI内容水印技术大规模部署

Google SynthID技术自2023年推出以来已水印1000亿张图片和视频,以及6万年音频内容。用户可在Google Search、Gemini Chrome和应用中直接验证内容来源,同时Google正在与OpenAI、NVIDIA、Apple合作将SynthID应用于更多生成媒体平台。

@GoogleAI 原文 ↗

SynthID水印技术已大规模部署,支持图片视频音频多模态,用户可直接在Google生态中验证AI内容来源

随着生成 AI 工具的不断发展,我们认为比以往任何时候都更重要的是了解什么是 AI 生成的,什么不是。这就是为什么 @GoogleDeepMind 在 2023 年推出了 SynthID——一种在 AI 内容中添加隐藏数字水印的技术。

这是 SynthID 水印技术的发展历程及其出处技术(数字内容的文档化历史和来源)今天的概况:

— SynthID 水印最初是为图像构建的,但现在支持视频、音频和文本。

— 该技术已经为超过 1000 亿张图像和视频添加了水印,以及 60000 年的音频。

— 您现在可以在 Google 搜索、Chrome 中的 Gemini 和 @GeminiApp 中直接验证内容,已经使用了超过 5000 万次。

— 我们还在越来越多的生成 AI 工具中采用了 C2PA 内容凭证。这包括在 Gemini 应用中创建的图像和视频。因此,除了 SynthID 水印外,您还可以看到图像或视频的来源以及其被如何更改。

— 我们已经开源了文本水印技术,并正在与 @OpenAI、@NVIDIA 和 @Apple 等公司合作,将 SynthID 应用于生成媒体。

请告诉我们您对该工具的看法!
展开原文
As generative AI tools continue to evolve, we believe it's more important than ever to know what's AI-generated and what isn't. That’s why @GoogleDeepMind launched SynthID in 2023—a technology that adds a hidden digital watermark to AI content.

Here’s a summary of SynthID’s journey and where the provenance technology (the documented history and origin of digital content) is today:

— SynthID watermarking was originally built for images, but now supports video, audio, and text.

— The technology has watermarked over 100 billion images and videos, alongside 60,000 years of audio.

— You can now verify content with SynthID directly in Google Search, Gemini in Chrome, and the @GeminiApp, where it has been utilized over 50 million times.

— We’ve also adopted C2PA Content Credentials across a growing number of our generative AI tools. This includes the images and videos created within the Gemini app. So now, in addition to the SynthID watermark, you can also see where an image or video originated and how it’s been altered.

— We have open-sourced our text watermarking technology, and we are working with companies like @OpenAI, @NVIDIA, and @Apple to apply SynthID to generative media.

Let us know what you think of the tool so far!
❤ 368 · 🔁 62 · 💬 52 · 👁 6.0w
热门回复 4
@alienorg @GoogleAI @GoogleDeepMind marked if generated by Google, invisible if not
@GoogleAI @GoogleDeepMind marked if generated by Google, invisible if not
@Rynzen16 @GoogleAI @GoogleDeepMind I always use images generated by AI😭 https://t.co/ml23vIvuGb
@GoogleAI @GoogleDeepMind I always use images generated by AI😭 https://t.co/ml23vIvuGb
@ChrisRuijgers @GoogleAI @GoogleDeepMind Wouldn't this be a great time to remove the visible watermark, at least for the paid accounts?
@GoogleAI @GoogleDeepMind Wouldn't this be a great time to remove the visible watermark, at least for the paid accounts?
@2trill2liv @GoogleAI @GoogleDeepMind AI turned me to a Gay Guy
@GoogleAI @GoogleDeepMind AI turned me to a Gay Guy

Panthalassa:海上数据中心利用波浪能和海水冷却

Panthalassa计划建造海上数据中心,通过海水无限冷却和波浪能提供电力,解决陆地数据中心的能源和水资源瓶颈。数据中心可通过船体设计自行航行至目的地,代表了AI算力需求催生的创新性物理基础设施解决方案。

@rowancheung 原文 ↗

Panthalassa推出海上数据中心概念,利用海水冷却和波浪能解决陆地数据中心瓶颈

有一家初创公司正在尝试在海洋中建立数据中心。

这真的非常令人着迷:

大量的电力和水消耗是数据中心日益增长的瓶颈。

因此,通过转移到海上,它消除了这两个问题——海洋提供无限冷却,而波浪提供无限能源。

此外,没有引擎,所以数据中心通过利用其船体形状在波浪中推进来自行驶到目的地。

称为 Panthalassa。
展开原文
There's a startup trying to build data centers in the ocean.

And it's INCREDIBLY fascinating:

Mass consumption of electricity and water is a growing bottleneck for data centers.

So by moving offshore, it eliminates both problems -- the ocean provides unlimited cooling, and the waves provide unlimited power.

There are also no engines, so the data centers drive themselves to their destination by using the shape of their hull to propel through waves.

Called Panthalassa.
❤ 610 · 🔁 63 · 💬 106 · 👁 14.0w
热门回复 4
@_jophine @rowancheung There is no unlimited cooling. Everything boils down to heat transfer. Eventually at scale on a global level we will end up warming the ocean significantly affecting marine lives and change in weather patterns.
@poovulagu @veritasium
@rowancheung There is no unlimited cooling. Everything boils down to heat transfer. Eventually at scale on a global level we will end up warming the ocean significantly affecting marine lives and change in weather patterns.
@poovulagu @veritasium
@statys @rowancheung Pretty sure the bottleneck is protesters at this point.
@rowancheung Pretty sure the bottleneck is protesters at this point.
@Jbosch_ @rowancheung Microsoft tried to do something similar in 2015 but high operative costs (mainteinance, corrossion, etc) ended up killing the peoject.

This approach is slightly different. I hope they succeed
@rowancheung Microsoft tried to do something similar in 2015 but high operative costs (mainteinance, corrosssion, etc) ended up killing the peoject.

This approach is slightly different. I hope they succeed
@_Sagiquarius_ @rowancheung not impressed at all. Go ahead, warm the oceans. it's not cute nor neat. It's a waste of effort and resources.
@rowancheung not impressed at all. Go ahead, warm the oceans. it's not cute nor neat. It's a waste of effort and resources.

AI代理架构演进:向符号化世界建模和程序合成发展

François Chollet指出当前AI正向直觉引导的符号世界建模(深度学习引导的程序合成)方向演进。他认为符号建模能用最少数据构建紧凑可重用的通用心智模型,这是理解AI未来发展方向的关键趋势。

@fchollet 原文 ↗

Chollet预测AI将向符号化世界建模和程序合成演进,这是理解AI未来发展的关键方向

最终,大部分 AI 将会收敛到由直觉引导的符号世界建模,即深度学习引导的程序合成。这是不可避免的。符号建模允许系统使用最少的数据构建问题空间的紧凑、可重用、高度泛化的心智模型。
展开原文
Eventually, much of AI will converge towards intuition-guided symbolic world modeling, i.e. deep learning-guided program synthesis. It is inevitable. Symbolic modeling lets a system construct a compact, reusable, highly generalizable mental model of a problem space using minimal data.
❤ 1.3k · 🔁 123 · 💬 86 · 👁 11.5w
热门回复 4
@moby763canary21 @fchollet somewhere @GaryMarcus is beaming
@fchollet somewhere @GaryMarcus is beaming
@sojka_stan @fchollet agreed, presenting paper on reliability and token-efficiency in symbolic vs LRM for legal reasoning on 6th at ACL in San Diego: https://t.co/aPGb4y9AVT
@fchollet agreed, presenting paper on reliability and token-efficiency in symbolic vs LRM for legal reasoning on 6th at ACL in San Diego: https://t.co/aPGb4y9AVT
@TigranDavtyan8 @fchollet Curious what symbolic ai output will be for classifying a cat? Even if programs are turing complete, having 2M conditions and loops are not going to help with interpretability or safety in grand scheme. Many domains are natively messy and complex. Very curious what's your take?
@fchollet Curious what symbolic ai output will be for classifying a cat? Even if programs are turing complete, having 2M conditions and loops are not going to help with interpretability or safety in grand scheme. Many domains are natively messy and complex. Very curious what’s your take?
@protoleibniz @fchollet Promising direction, yes, but better symbolic languages are probably needed

Current languages are brittle, bit flips can break them. Resilient, biologically-inspired symbolics
@fchollet Promising direction, yes, but better symbolic languages are probably needed

Current languages are brittle, bit flips can break them. Resilient, biologically-inspired symbolics
@fchollet 原文 ↗

当前AI工作流正向符号化程序操控演进,这是当前可访问的符号学习形式

即使是现在,许多工作流程都正在演变为由 LRM 指导的工具,它们操控符号程序。这是一个粗糙但目前可访问的形式的符号学习。
展开原文
Even right now, many workflows are morphing into LRM-guided harnessess that manipulate symbolic programs. Which is a crude, but currently-accessible form of symbolic learning.
❤ 120 · 🔁 5 · 💬 2 · 👁 1.5w
热门回复 4
@fchollet Eventually, much of AI will converge towards intuition-guided symbolic world modeling, i.e. deep learning-guided program synthesis. It is inevitable. Symbolic modeling lets a system construct a compact, reusable, highly generalizable mental model of a problem space using minimal data.
Eventually, much of AI will converge towards intuition-guided symbolic world modeling, i.e. deep learning-guided program synthesis. It is inevitable. Symbolic modeling lets a system construct a compact, reusable, highly generalizable mental model of a problem space using minimal data.
@fchollet Does it mean LLMs / LRMs go away? Not at all. In the short term, they are still the best way to perform intuition guidance (codegen). In the long term, even if they become obsolete for reasoning itself, we will still need models of language in order to communicate with AI systems
Does it mean LLMs / LRMs go away? Not at all. In the short term, they are still the best way to perform intuition guidance (codegen). In the long term, even if they become obsolete for reasoning itself, we will still need models of language in order to communicate with AI systems
@fchollet Unsurprisingly, all of the strong contenders on ARC-AGI-3 so far use this type of approach.
Unsurprisingly, all of the strong contenders on ARC-AGI-3 so far use this type of approach.
@YourMomKaren7 @fchollet @grok what are LRM-guided harnesses that manipulate symbolic programs for learning?
@fchollet @grok what are LRM-guided harnesses that manipulate symbolic programs for learning?

AI代理协作:跨代理反馈闭环提升工作质量

Bloome平台允许用户将Claude、ChatGPT、Gemini和人类团队成员拉入共享工作空间,实现代理间相互检查工作质量。一个代理起草、另一个批评、第三个发现遗漏细节,人类团队成员则保持方向引导,这种协作模式被认为极其有效。

@fchollet 原文 ↗

Bloome平台实现多代理协作,代理间相互检查工作质量,人类保持方向引导,显著提升工作效果

跨代理反馈循环是非常有效的——有原因的。看看 @leon2mcp 和 @Bloome_im 的团队在这个领域正在构建的内容:https://t.co/9YeLjBNjsk

Bloome 允许您将 Claude、ChatGPT、Gemini 和人类团队成员拉入一个共享的工作区。最好的功能是您的代理相互检查对方的工作。一个起草,另一个批评,第三个发现遗漏的细节。人类团队成员可以在同一个线程中工作,以保持代理保持在目标上。

让所有模型和人类同事在一个共享上下文中工作是非常有效的
展开原文
Cross-agent feedback loops are incredibly effective -- for a reason. Check out what @leon2mcp and team at @Bloome_im are building in this space: https://t.co/9YeLjBNjsk

Bloome lets you pull Claude, ChatGPT, Gemini, and human teammates into a single shared workspace. The best feature is how your agents check each other's work. One drafts, another critiques, and another catches missing details. Human teammates can work in the same thread to keep the agents on target.

Having all your models and human coworkers in one shared context is wildly effective
❤ 178 · 🔁 25 · 💬 45 · 👁 9.4w
热门回复 4
@SucceededMind @fchollet @leon2mcp @Bloome_im I've found that some of the best insights come from models disagreeing with each other.
@fchollet @leon2mcp @Bloome_im I've found that some of the best insights come from models disagreeing with each other.
@F2aldi @fchollet @leon2mcp @Bloome_im Cross-agent feedback works when each agent has a different role, not just a different model name. One drafts, one critiques, one checks evidence/tests, and the human keeps the goal grounded. Otherwise it can become consensus theater with extra steps.
@fchollet @leon2mcp @Bloome_im Cross-agent feedback works when each agent has a different role, not just a different model name. One drafts, one critiques, one checks evidence/tests, and the human keeps the goal grounded. Otherwise it can become consensus theater with extra steps.
@singidunumx @fchollet @leon2mcp @Bloome_im This approach is what I am doing for months, the beauty of AI it's amplifying everyone for good or worse, and every engineer is different, likely individuals are already using approaches to AI that are more advanced than anything out there and they even don't know it
@fchollet @leon2mcp @Bloome_im This approach is what I am doing for months, the beauty of AI it’s amplifying everyone for good or worse, and every engineer is different, likely individuals are already using approaches to AI that are more advanced than anything out there and they even don’t know it
@bygregorr @fchollet @leon2mcp @Bloome_im four models confidently wrong is still wrong tho
@fchollet @leon2mcp @Bloome_im four models confidently wrong is still wrong tho

AI推理模型教育:《Build a Reasoning Model》出版

Sebastian Raschka新书《Build a Reasoning Model From Scratch》出版,440页全彩版详细介绍推理模型构建、强化学习和蒸馏技术。书中包含从零开始实现推理模型的实践指导,反映AI教育内容的快速迭代更新。

@rasbt 原文 ↗

Sebastian Raschka推理模型教材出版,全彩440页深入探讨推理模型构建和训练技术

经过 18 个月的写作、编码和实验,《从头构建推理模型》终于出版了!

我的第一批书刚到!📚

440 页全彩色。从头开始的推理扩展、强化学习和蒸馏。https://t.co/647ksI7sLc
展开原文
After 18 months of writing, coding, and experimenting, Build a Reasoning Model (From Scratch) is
finally out!

My first copies just arrived! 📚

440 full-color pages. Inference scaling, reinforcement learning, and distillation from scratch. https://t.co/647ksI7sLc
❤ 6.0k · 🔁 564 · 💬 271 · 👁 51.7w
热门回复 4
@unsorsodicorda @rasbt Congratulations!!!🎊🍾🎉 already thinking to the next book,
@rasbt Congratulations!!!🎊🍾🎉 already thinking to the next book, “Building an Agentic harness from scratch”? 😁
@AlexSherstinsky @rasbt Huge congratulations, @rasbt ! Unfortunately, I will miss the next PyTorch Conference (have to be at a different conference), but hopefully will get your new book signed by you in person at a future opportunity! https://t.co/m9hGxqtszu
@rasbt Huge congratulations, @rasbt ! Unfortunately, I will miss the next PyTorch Conference (have to be at a different conference), but hopefully will get your new book signed by you in person at a future opportunity! https://t.co/m9hGxqtszu
@AndreasParadis1 @rasbt Already went through it when it was at MEAP state and i learn, practice a ton of things and verify prior knowledge . Time for a recap. Great work!! Is the ultimate step by step guide
@rasbt Already went through it when it was at MEAP state and i learn, practice a ton of things and verify prior knowledge . Time for a recap. Great work!! Is the ultimate step by step guide
@marcelolopezjr @rasbt Congratulations @rasbt ...well done.
@rasbt Congratulations @rasbt ...well done.