每日自动抓取:arXiv 当日新论文、近半年高被引论文、Hacker News 热门。仅供学习参考。
arXiv 今日新论文(AI / 机器学习 / 量化金融)
1. Objective vs. Search: Decomposing What Makes a Good Tokeniser
Ahmetcan Yavuz, Clara Meister, Tiago Pimentel 等 · 2026-09-16
Two dominant tokenisation algorithms are used by modern language models: byte-pair encoding (BPE) and UnigramLM. These differ along two orthogonal axes: their optimisation objective (compression vs. log-likelihood) and …
2. A Zeroth-Order Paradigm for LLM Preference Alignment
Peter Chen, Xi Chen, Wotao Yin 等 · 2026-09-16
Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates …
3. PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
Sara Pieri, Evangelos Kazakos, Shizhe Chen 等 · 2026-09-16
Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded. Current vision-language models (VLMs) can generate fluent and detailed image captions, but …
4. Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation
Guanhua Ji, Tianyu Li, Dayoon Suh 等 · 2026-09-16
Recent advances in video generation allow robots to learn manipulation trajectories from generated videos. However, these approaches produce purely kinematic trajectories that lack force information, causing failures in …
5. ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments
Hejia Geng, Zesen Huang, Haoyang Li 等 · 2026-09-16
Scientific code repositories encode decades of human knowledge in executable models, methods, and tools. Yet fragmented toolchains, implicit domain conventions, and specialized correctness criteria make this knowledge …
6. Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments
João Meneses dos Santos, Arlindo L. Oliveira 等 · 2026-09-16
Language agents remain brittle in interactive environments, where success requires long-horizon state tracking, valid action execution, and recovery from failed steps. We extend SwiftSage, a dual-process agent that …
7. Affora: A Design System for Agent-Friendly Interfaces
Jin Gao 等 · 2026-09-16
Computer-use agents increasingly operate software designed for people, but interfaces often leave actions or task state unclear to machine readers. We present Affora, a design system that supports both readers while …
8. Flag Game: A Toy Model for Mechanistic Swarm Interpretability
Elizabeth Pavlova, Hidenori Tanaka 等 · 2026-09-16
Emergent coordinated behaviors of AI agents are starting to present critical safety risks. A key phenomenon driving these behaviors is the rapid formation and spread of beliefs about the world, and mechanistic …
9. Playing log(N)-Questions over Wikipedia Abstracts: Communication Efficiency Between Paired Frontier Models
Peter Potash 等 · 2026-09-16
We evaluate six frontier language models on the two-agent $\log(N)$-Questions game. A questioner sees $N$ Wikipedia lead paragraphs and must identify a secretly chosen target using exactly $\log_2 N$ yes/no questions. …
10. rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference
Kaijun Zhou, Zhiyang Li, Le Chen 等 · 2026-09-16
Factory work is a promising early scenario for embodied AI: assigning repetitive manual jobs to robots has clear economic payoff, and a structured station keeps the jobs tractable for current policies. …
11. Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
Leon Bergen, Usha Bhalla, Andrew Lee 等 · 2026-09-16
As models scale, reward hacking becomes more frequent, more sophisticated, and more consequential. Does it leave a telltale signature in model representations? This work analyzes how reward hacking is represented …
12. Prepared Or Unprepared? Evaluating Healthcare Workforce Readiness for Clinical Adoption of Artificial Intelligence in Nigeria
Abbas M. Rabiu, Abdulrazaq A. Zubair, Um-mulkhairi Ibrahim 等 · 2026-09-16
Artificial intelligence (AI) is increasingly integrated into healthcare systems worldwide, yet its successful clinical adoption depends critically on workforce readiness, particularly in low- and middle-income countries …
13. Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation
Daniel P. Jeong, Charles Q. Li, Hossein Hosseiny 等 · 2026-09-16
Radiologists follow heterogeneous reporting practices. Two radiologists examining the same image and identifying the same clinical findings might nevertheless compose superficially distinct reports, varying in …
14. Securing quantum error correction against misleading advice from AI agents
A. Barış Özgüler 等 · 2026-09-16
Can an attacker turn influence over an artificial intelligence (AI) adviser into a harmful quantum error-correction update? We identify an ambiguity in passive syndrome records that obstructs recovery selection, then …
15. MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education
Luyao Zhu, Xun Wei Yee, Wei Li 等 · 2026-09-16
Large vision-language models have achieved remarkable progress in multi-modal understanding, yet their capabilities in educational settings remain insufficiently evaluated. In AI-assisted language learning, models must …
共 15 篇。
近半年高被引论文
AI / 机器学习方向
量化金融方向
Hacker News 热门(科技 / 创业 / 投资风向)
| 热度 | 讨论 | 标题 |
|---|---|---|
| 2170 赞 | 243 评 | Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations |
| 705 赞 | 287 评 | Nvidia announces native GPU programming in Rust |
| 559 赞 | 118 评 | Training a 4B model to produce 81% faster query plans than Postgres |
| 528 赞 | 239 评 | Small programming tricks |
| 440 赞 | 116 评 | Xiaomi Mimo 2.6 live post-training dashboard |
| 410 赞 | 342 评 | AWS says it can't restore some data from mideast facilities struck by Iran |
| 295 赞 | 64 评 | Performance Improvements in .NET 11 |
| 245 赞 | 147 评 | Backups Aren't Simple |
| 211 赞 | 34 评 | Reversing Factorio's RNG |
| 209 赞 | 79 评 | The engineering behind the US Strategic Petroleum Reserve |
| 209 赞 | 33 评 | Breaking the 1.58-bit Barrier for Ternary LLMs |
| 194 赞 | 79 评 | Japan's book scene is moving from bookstores to libraries |
| 178 赞 | 65 评 | Keys Not Included: recovering the signing keys for US driver's license barcodes |
| 152 赞 | 234 评 | Anecdotally, programmers dislike "reduce" |
| 139 赞 | 63 评 | OpenSpec – A lightweight and configurable AI spec framework |
本日报由服务器每日自动抓取生成于凌晨,原文链接均已附在上。