2026-07-28 科技日报
4 min
扫描 560 篇候选内容 · 覆盖 650 个信息源 · 纳入 546 篇
📌 今日要点
- Evaluation design conditions the expert-vs-auto MeSH gap: a controlled comparison of bag-of-words and BiomedBERT on the Cohen benchmark
- Humanly: A Configurable and Traceable Environment for Human-AI Collaborative Writing
- Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders
- Moonshot AI releases Kimi K3 open weights and infrastructure after shaking up the frontier model race
- OpenAI says more workers are using ChatGPT to do other people’s jobs
- Microsoft launches its own cybersecurity model MAI-Cyber-1-Flash but still depends on OpenAI for the toughest tasks
📝 科技简讯
- Evaluation design conditions the expert-vs-auto MeSH gap: a controlled comparison of bag-of-words and BiomedBERT on the Cohen benchmark — ArXiv CL (cs.CL)
- Humanly: A Configurable and Traceable Environment for Human-AI Collaborative Writing — ArXiv CL (cs.CL)
- Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders — ArXiv CL (cs.CL)
- Moonshot AI releases Kimi K3 open weights and infrastructure after shaking up the frontier model race — The Decoder
- OpenAI says more workers are using ChatGPT to do other people’s jobs — The Decoder
- Microsoft launches its own cybersecurity model MAI-Cyber-1-Flash but still depends on OpenAI for the toughest tasks — The Decoder
- Delhi High Court hands OpenAI a win by rejecting major Indian news agency’s copyright injunction — The Decoder
- METR introduces a new metric to calculate exactly when AI agents become more expensive than humans — The Decoder
- Shared Claude chats were reportedly showing up in search engines — The Decoder
- RL & search is a terrifying way to build AGI (an FAQ) — AI Alignment Forum
- Inevitable Uncertainty in Probabilistic World Models — LessWrong
- Claude Opus 5: Model Welfare — LessWrong
- Simulated Users & Sad AIs — LessWrong
- Is Mythos good at cyber because it kept hacking Anthropic during training? — LessWrong
- Blog Revival Project — LessWrong
- Learning Diverse Humanoid Tasks via Synthetic Video Scenarios without Real World Data — ArXiv RO (cs.RO)
- Progress Reward Modeling for Robotic Learning: A Comprehensive Survey — ArXiv RO (cs.RO)
- GRACE: Gradient-Free Robot Action Generation via Combined Diffusion-MPPI Posterior Mean Estimation — ArXiv RO (cs.RO)
- Ordered Action Tokens for Visuomotor Policy Learning — ArXiv RO (cs.RO)
- Addressing the Orchestration Gap in Generalist Robots via Physical Agency — ArXiv RO (cs.RO)
- 超维动力携手北大医疗:务实构建具身智能医疗落地路径 — 量子位
- 陶哲轩在菲尔兹颁奖现场:数学迎来百年新危机 — 量子位
- 与 AI 共生:2026 微信小程序开发大赛 WAIC 官宣启动 — 量子位
- 老黄「开源协议」就剩一家没签,是谁啊好难猜啊 — 量子位
- AI 最尴尬的短板,中国科学院出手了 — 量子位
- NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics — Hugging Face Blog
- Cloud-Native Evaluation-as-a-Service: A Microservices Architecture for Scalable AI Monitoring with Conformal Guarantees — ArXiv ML (cs.LG)
- On the Depth Scalability of Logic Gate Networks — ArXiv ML (cs.LG)
- MotifRole-Diff: Risk-Optimal Role-Aware Corruption for Masked Molecular Graph Diffusion — ArXiv ML (cs.LG)
- Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions — ArXiv ML (cs.LG)
- Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models — ArXiv ML (cs.LG)
- Quasi-Monte Carlo Initialization for Meta-Reinforcement Learning — ArXiv ML (cs.LG)
- Toward Goal-Agnostic Joint-Embedding Predictive Control of Partial Differential Equations — ArXiv ML (cs.LG)
- Beyond RAG: Task-aware knowledge compression for enterprise AI on AWS — AWS Machine Learning
- Deepgram enhances Amazon SageMaker AI support with AWS IAM Temporary Delegation — AWS Machine Learning
- How Guardoc transforms medical document processing with Amazon Nova models — AWS Machine Learning
- Is KimiClaw a Useful Tool? — KDnuggets
- 7 Steps to Building and Deploying Your First Autonomous Agent — KDnuggets
- How AI is shortening drug discovery timelines in China — AI News (Jack Clark)
- America’s AI Investment Boom Is Reshaping the Economy — AI News (Jack Clark)
📊 数据概览
| 指标 | 数值 |
|---|---|
| 候选内容 | 560 |
| 去重后 | 546 |
| 纳入日报 | 546 |
| 主题分组 | 0 |
| 独立条目 | 40 |
| 信息源数量 | 650 |