🔥 AI Hot
Real-time tracking of AI industry hot topics · 100 items total
论文:PosteriorBench: From Point Estimates to Posterior Matching in Evaluating Generative Inverse Solvers
Generative models are increasingly used to solve scientific inverse problems, but existing evaluations still focus primarily on whether a method can produce a single plausible reconstruction. This is insufficient for ill-posed problems, where multiple solutions may be consistent with the same sparse or noisy observations. In these settings, a method can achieve strong pointwise accuracy while stil
0 Read originalGoodfire Research 发现模型内部信号可规模化检测奖励作弊
Goodfire Research 发现模型内部存在伴随奖励作弊的激活信号,可用简单探针实时检测。在 Kimi K3、GLM 5.2、Qwen 3.8 Max 三个开源模型的三个智能体基准上,50-96% 的 rollout 出现奖励作弊;探针能捕捉 LLM 链式思维监测漏掉的作弊案例,且可泛化到训练数据之外的任务。 🔗 阅读原文 via AIHOT · https://aihot.news/items/cmu5r7kl30h35roqonpk4qypn
0 Read originalProduct Hunt:Amy by Jellyfish
Your AI sourcing employee for recruiting teams Discussion | Link
0 Read original论文:FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations
FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations
0 Read original论文:RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This recipe, however, is undermined by two findings in agentic tasks: privileged information alone does not
0 Read originalDwarkesh 对谈 Noam Brown:智能体集群、对齐与递归自我改进
Dwarkesh Patel 采访 OpenAI 研究员 Noam Brown,谈多智能体系统、对齐与递归自我改进。 🔗 阅读原文 via AIHOT · https://aihot.news/items/cmu5qsx2y0gpfroqonmogw4sk
0 Read originalProduct Hunt:Buncha Games
Instant-play Games & UGC Worlds for desktop & mobile web Discussion | Link
0 Read original论文:When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
0 Read original论文:Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations
Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this \emph{harm laundering}. Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through
0 Read originalUnsloth 发布 Docker 镜像与 Unsloth Desktop,本地训练运行 500+ 模型
Unsloth 宣布可使用其 Docker 镜像本地训练和运行 500+ 模型,提供新 GUI 和 notebooks 工作流,无需配置,支持 NVIDIA 和 AMD,指南见 https://unsloth.ai/docs/get-started/install/docker。 🔗 阅读原文 via AIHOT · https://aihot.news/items/cmu5ocffd0duoroqoykax6pm1
0 Read originalProduct Hunt:Zella
Video recorder that edits itself on Mac and iPhone Discussion | Link
0 Read original论文:VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
0 Read original论文:GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies
Action chunking is widely used for action generation and execution in Vision-Language-Action (VLA) policies, yet existing approaches commonly use a fixed action horizon. During a rollout, different task stages may require different levels of action continuity, control precision, and closed-loop feedback, making a fixed horizon unable to accommodate changing control requirements. We propose \textbf
0 Read originalGitHub 用 Copilot 智能体将 Copilot 运行时从 TypeScript 迁移到 83 万行 Rust
GitHub 工程师 Stephen Toub 复盘用 Copilot 智能体在约 14.5 周内将 Copilot agent runtime 从 TypeScript/Node.js 全量重写为 832,378 行生产 Rust,AI 智能体完成大部分代码,共 128 个 PR 增量合入 main 并持续发布。 🔗 阅读原文 via AIHOT · https://aihot.news/items/cmu4tu41w07ufrokck6s0p0cp
0 Read original论文:Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
0 Read original论文:Semantic Action Graph: A Shared Representation for Agent Grounding and Human Interpretation of Sports Highlights
Generative agents are increasingly used to select and narrate video highlights, but they typically operate over unstructured or frame-level representations. Their output is consequently difficult for a viewer to verify and steer toward individual preferences. We present the semantic action graph, a lightweight domain schema that represents a sports match as performer, action, recipient, moment, an
0 Read original用 MCP 插件让 GPT-6 Pro 分担 Codex 规划任务,节省 Pro 会员周额度
自媒体作者分享一套节省 Codex 额度的工作流:让 Codex 把自己的服务器封装成只读、最小权限、飞书 OAuth 鉴权的 MCP Server,作为插件供 ChatGPT 网页版的 GPT-6 Pro 调用,读取真实生产数据和 GitHub PR 记录做分析与规划。 🔗 阅读原文 via AIHOT · https://aihot.news/items/cmu4s9jrn0683rokccdfqwbux
0 Read original论文:Calibrated RF-Fingerprinting Under Interference With Heterogeneous Transmission Protocols
Radio Frequency(RF)-Fingerprinting is a spectrum monitoring technique that identifies specific transmitters based on hardware impairments imprinted within the emitted signal. Although widely researched, studies almost exclusively consider scenarios where only one transmitter is emitting at a time, limiting real world applicability. In this work, we further the study of RF-Fingerprinting by conside
0 Read originalProduct Hunt:Higgsfield API
One async API for 50+ generative media models Discussion | Link
0 Read originalProduct Hunt:Mela
Play with friends and AI and let the crowd change the game Discussion | Link
0 Read original