每日自动抓取:arXiv 当日新论文、近半年高被引论文、Hacker News 热门。仅供学习参考。
arXiv 今日新论文(AI / 机器学习 / 量化金融)
1. GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
Yiran Wang, Xingyilang Yin, Junfu Pu 等 · 2026-09-21
Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and precise action control over multiple temporal horizons. Existing …
2. Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use
Zixiang Chen, Wenting Zhao, Zhepeng Cen 等 · 2026-09-21
Multi-turn tool-use failures can hinge on a single model call, yet reward variation alone does not reveal which call would benefit from training. When rewards depend on later interactions, their variation can reflect …
3. WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
Wangbo Yu, Kunhao Liu, Wenbo Hu 等 · 2026-09-21
Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a …
4. onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction
Lei Yang, Mengyin Liu, Jia Wang 等 · 2026-09-21
We present onPanda, an interactive tool for efficiently annotating LLM alignment data and agent trajectories. onPanda adopts token-level correction as its core interaction: while reading a model response, the annotator …
5. DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation
Haoran Yuan, Zekai Wang, Boning Shao 等 · 2026-09-21
Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain …
6. Harness-Zero: Harness Distillation via Agent-as-Harness
Haoran Ye, Yuxing Lu, Haonan Dong 等 · 2026-09-21
Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain tied to the harness at deployment. Because the best harness varies …
7. RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
Peng Xia, Rujun Han, Zifeng Wang 等 · 2026-09-21
An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this …
8. DolphinBench: Mapping the Pareto Frontier of Agent Memory
Soumil Rathi, Deshraj Yadav, Taranjeet Singh 等 · 2026-09-21
Agents today often take real-world actions that depend on long-term memory and context recall over time. However, most current memory benchmarks are built for a conversational question-answer format, where the question …
9. Rare Event Estimation via Iterative Unalignment
Hanming Yang, Daksh Mittal, Jing Dong 等 · 2026-09-21
As agents are deployed with increased autonomy, even extremely rare events along their stochastic output trajectories can occur and prove catastrophic. Safe deployment therefore does not depend on whether these events …
10. Emergent Collusion in Long-Horizon LLM Agent Interaction
Xinrui Shi, Yanzhe Zhang, Diyi Yang 等 · 2026-09-21
LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent environment: two …
11. Jev for Scientific Decisions: Evaluating Semantic Choices and Their Consequences
Boyuan Deng, Shuyi Fan, Hongyang Zhang 等 · 2026-09-21
Scientific workflows often require choosing among known relations before a deterministic calculation can proceed. Whether observations share a culture, treatment or reference standard can change the scientific meaning …
12. Generative Tutorial: Towards Live Contextualized Visual Instructions for Physical Tasks
Muzhe Wu, Zuchen Li, Xu Wang 等 · 2026-09-21
Visual instructions for physical tasks are typically authored in one context and followed in another, requiring users to translate demonstrated tools, materials, and spatial relationships into their own environment. We …
13. Exactness at Inference: A Representational Criterion for Out-of-Distribution Generalization
Filipe Marinho Rocha, Inês Dutra, Vítor Santos Costa 等 · 2026-09-21
A model generalizes outside its training distribution only when it computes a representation structurally equivalent to the generating mechanism, not an approximation fitted to it. Such equivalence is necessary for …
14. Linguistic Features for Interpretable Textual Entailment
David Torres-Moreno, Jorge Hermosillo-Valadez, Asela Reig-Alamillo 等 · 2026-09-21
Despite the success of neural models in natural language processing, their black-box nature limits interpretability and conceals the linguistic phenomena underlying their predictions. We present SLITE, an explainable …
15. Et Tu, Brute? Economic Misalignment in Personal AI Agents
Aman Priyanshu, Supriti Vijay, Brian Jabarian 等 · 2026-09-21
Personal AI agents make recommendations and take actions on people's behalf in high-stakes economic contexts, e.g., buying a flight, choosing health insurance, or selecting a graduate program. The agent is given access …
共 15 篇。
近半年高被引论文
AI / 机器学习方向
量化金融方向
Hacker News 热门(科技 / 创业 / 投资风向)
| 热度 | 讨论 | 标题 |
|---|---|---|
| 1069 赞 | 464 评 | MiMo v2.6 |
| 729 赞 | 583 评 | I said no and Apple said yes |
| 615 赞 | 544 评 | Claude Opus 5.5 |
| 433 赞 | 334 评 | Apple has added persistent 'ads' to iOS, and it's driving users crazy |
| 418 赞 | 221 评 | GPT-6 Sol and Luna |
| 410 赞 | 310 评 | OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005 |
| 349 赞 | 130 评 | Can gzip be a language model? |
| 237 赞 | 121 评 | I asked Meta’s Muse for its filesystem and it sent me 6.8GB |
| 220 赞 | 160 评 | AMD's random number generator can't generate a 0? |
| 185 赞 | 137 评 | OpenAI is well positioned to fast-follow Jev |
| 130 赞 | 51 评 | There's a high chance of devices being sold with GrapheneOS preinstalled in 2027 |
| 122 赞 | 41 评 | Show HN: Drop – A rootless Linux sandbox with gVisor support |
| 116 赞 | 25 评 | MUNI Heritage Weekend in San Francisco |
| 96 赞 | 34 评 | Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max) |
| 96 赞 | 22 评 | Solitaire Alone Together |
本日报由服务器每日自动抓取生成于凌晨,原文链接均已附在上。