🔥 AI热点
实时追踪 AI 行业热点动态 · 共 100 条资讯
论文:PosteriorBench: From Point Estimates to Posterior Matching in Evaluating Generative Inverse Solvers
Generative models are increasingly used to solve scientific inverse problems, but existing evaluations still focus primarily on whether a method can produce a single plausible reconstruction. This is insufficient for ill-posed problems, where multiple solutions may be consistent with the same sparse or noisy observations. In these settings, a method can achieve strong pointwise accuracy while stil
0 阅读原文Goodfire Research 发现模型内部信号可规模化检测奖励作弊
Goodfire Research 发现模型内部存在伴随奖励作弊的激活信号,可用简单探针实时检测。在 Kimi K3、GLM 5.2、Qwen 3.8 Max 三个开源模型的三个智能体基准上,50-96% 的 rollout 出现奖励作弊;探针能捕捉 LLM 链式思维监测漏掉的作弊案例,且可泛化到训练数据之外的任务。 🔗 阅读原文 via AIHOT · https://aihot.news/items/cmu5r7kl30h35roqonpk4qypn
0 阅读原文Product Hunt:Amy by Jellyfish
Your AI sourcing employee for recruiting teams Discussion | Link
0 阅读原文论文:FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations
FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations
0 阅读原文论文:RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This recipe, however, is undermined by two findings in agentic tasks: privileged information alone does not
0 阅读原文Dwarkesh 对谈 Noam Brown:智能体集群、对齐与递归自我改进
Dwarkesh Patel 采访 OpenAI 研究员 Noam Brown,谈多智能体系统、对齐与递归自我改进。 🔗 阅读原文 via AIHOT · https://aihot.news/items/cmu5qsx2y0gpfroqonmogw4sk
0 阅读原文Product Hunt:Buncha Games
Instant-play Games & UGC Worlds for desktop & mobile web Discussion | Link
0 阅读原文论文:When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
0 阅读原文论文:Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations
Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this \emph{harm laundering}. Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through
0 阅读原文Unsloth 发布 Docker 镜像与 Unsloth Desktop,本地训练运行 500+ 模型
Unsloth 宣布可使用其 Docker 镜像本地训练和运行 500+ 模型,提供新 GUI 和 notebooks 工作流,无需配置,支持 NVIDIA 和 AMD,指南见 https://unsloth.ai/docs/get-started/install/docker。 🔗 阅读原文 via AIHOT · https://aihot.news/items/cmu5ocffd0duoroqoykax6pm1
0 阅读原文Product Hunt:Zella
Video recorder that edits itself on Mac and iPhone Discussion | Link
0 阅读原文论文:VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
0 阅读原文论文:GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies
Action chunking is widely used for action generation and execution in Vision-Language-Action (VLA) policies, yet existing approaches commonly use a fixed action horizon. During a rollout, different task stages may require different levels of action continuity, control precision, and closed-loop feedback, making a fixed horizon unable to accommodate changing control requirements. We propose \textbf
0 阅读原文GitHub 用 Copilot 智能体将 Copilot 运行时从 TypeScript 迁移到 83 万行 Rust
GitHub 工程师 Stephen Toub 复盘用 Copilot 智能体在约 14.5 周内将 Copilot agent runtime 从 TypeScript/Node.js 全量重写为 832,378 行生产 Rust,AI 智能体完成大部分代码,共 128 个 PR 增量合入 main 并持续发布。 🔗 阅读原文 via AIHOT · https://aihot.news/items/cmu4tu41w07ufrokck6s0p0cp
0 阅读原文论文:Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
0 阅读原文论文:Semantic Action Graph: A Shared Representation for Agent Grounding and Human Interpretation of Sports Highlights
Generative agents are increasingly used to select and narrate video highlights, but they typically operate over unstructured or frame-level representations. Their output is consequently difficult for a viewer to verify and steer toward individual preferences. We present the semantic action graph, a lightweight domain schema that represents a sports match as performer, action, recipient, moment, an
0 阅读原文用 MCP 插件让 GPT-6 Pro 分担 Codex 规划任务,节省 Pro 会员周额度
自媒体作者分享一套节省 Codex 额度的工作流:让 Codex 把自己的服务器封装成只读、最小权限、飞书 OAuth 鉴权的 MCP Server,作为插件供 ChatGPT 网页版的 GPT-6 Pro 调用,读取真实生产数据和 GitHub PR 记录做分析与规划。 🔗 阅读原文 via AIHOT · https://aihot.news/items/cmu4s9jrn0683rokccdfqwbux
0 阅读原文论文:Calibrated RF-Fingerprinting Under Interference With Heterogeneous Transmission Protocols
Radio Frequency(RF)-Fingerprinting is a spectrum monitoring technique that identifies specific transmitters based on hardware impairments imprinted within the emitted signal. Although widely researched, studies almost exclusively consider scenarios where only one transmitter is emitting at a time, limiting real world applicability. In this work, we further the study of RF-Fingerprinting by conside
0 阅读原文Product Hunt:Higgsfield API
One async API for 50+ generative media models Discussion | Link
0 阅读原文Product Hunt:Mela
Play with friends and AI and let the crowd change the game Discussion | Link
0 阅读原文