前沿情报日报 · 2026-09-22

飘飘飘
发布于 2026-09-22 / 0 阅读
0

前沿情报日报 · 2026-09-22

每日自动抓取:arXiv 当日新论文、近半年高被引论文、Hacker News 热门。仅供学习参考。

arXiv 今日新论文(AI / 机器学习 / 量化金融)

1. Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
Hongyang Du, Lan Yan, Christian Flores 等 · 2026-09-18

Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle. We introduce a continual …

2. Cross-sector generalization of accident-process role classification in occupational accident narratives
Aho Yapi, Pierre Latouche, Arnaud Guillin 等 · 2026-09-18

Occupational accident narratives contain valuable information about work situations, unfavourable conditions, accident events, and their consequences. Automatically structuring these narratives can facilitate …

3. CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
Bowen Ye, Lei Li, Shicheng Li 等 · 2026-09-18

Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while existing methods typically rely on …

4. Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw
Renkai Ma, Ruyuan Wan, Xuan Lu 等 · 2026-09-18

Users increasingly delegate work to autonomous AI agents, yet evaluations typically measure task completion rather than the values users prioritize. Using Value Sensitive Design, we analyzed, with LLM assistance, 73,093 …

5. Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention
Andre Bacellar 等 · 2026-09-18

Multi-hop retrieval failures are not uniformly distributed across queries: they cluster in structurally predictable subpopulations. We prove two results formalizing this structure. First (CWAR Reducibility): …

6. An Interpretable Memory Decision Controller for LLM Agents Based on Three-Signal Complementarity: Decoupling Confidence and Consistency
Yiming Zhang, Jinghong Zhang, Haoran Zhao 等 · 2026-09-18

Memory systems for large language models have focused predominantly on efficient retrieval, whereas the decision of whether retrieved memories should be trusted has received comparatively little attention. When the …

7. Gricea: An Open Science Platform for Conversational AI Research
Nikhil Sharma, Yunlin Gong, Xinyang Cheng 等 · 2026-09-18

We need studies on conversational AI (CAI) at scale to understand human behavior and shape CAI design. However, fragmented reporting of systems and study configurations hinders replication, extension, and knowledge …

8. QuranicMMLU: A Cognitively-Aware Benchmark for Evaluating Generative AI Solutions on Quranic Linguistic Knowledge
Rawan El Ghali, Umm Kulsoom, Anas Madkoor 等 · 2026-09-18

We introduce QuranicMMLU, a benchmark for evaluating generative AI on Quranic Arabic across multiple dimensions of linguistic complexity. Existing Quranic benchmarks center on general question answering and semantic …

9. DiaVLo: Diagnosing Behaviours of Vision-Language Models
Lorenzo Corti, Jie Yang 等 · 2026-09-18

Vision-language models (VLMs) rely on storing and transferring appropriate information across their sub-components. Verifying that the VLMs exhibit desired behaviours, while avoiding harmful ones, is central to their …

10. Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention
Richard Zhe Wang 等 · 2026-09-18

Gating the value pathway of attention reportedly improves language model pretraining, and prior studies disagree on why. We argue and provide experimental evidence that such gates supply two different things that …

11. RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
Shuai Bai, Jiayong Deng, Yikun Fu 等 · 2026-09-18

Computer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development through code and the command line. Real digital work requires both, interleaved rather than stacked end …

12. Bayesian Belief Layer for Controllable Opinion Dynamics in LLM Agents
Hafsa Akbar, Daniel Platnick, Marjan Alirezaie 等 · 2026-09-18

LLM agents in social simulation revise their opinions implicitly, in context: how open an agent is to persuasion can neither be specified nor verified, and collective outcomes inherit the model's training prior. We …

13. A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal
Hiskias Dingeto 等 · 2026-09-18

Large language models can hold knowledge they do not report. A model may sandbag on a capability evaluation, or answer against what it internally knows, and its outputs alone cannot tell whether it is hiding an answer …

14. Moral Entropy: Auditing Bias and Uncertainty in Moral Judgment
Maciej Skorski 等 · 2026-09-18

Most work in computational ethics treats annotator disagreement on moral content as noise to be voted away, collapsed into majority vote or the more permissive any-annotator rule the moment a single annotator flags an …

15. NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
Jagadeesh Balam, Travis Bartley, Edresson Casanova 等 · 2026-09-18

We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combines a streaming speech encoder and decoder-only language model with …

共 15 篇。

近半年高被引论文

AI / 机器学习方向

论文 被引
A Survey of Large Language Models 1550 次
Polyendocrine metabolic ovarian syndrome, the new name for polycystic ovary syndrome: a multistep global consensus process 227 次
Sycophantic AI decreases prosocial intentions and promotes dependence 118 次
The MetroVolt Data-Center Burner: Direct-DC Campus Power, Q_E Closure, and the Plug Requirement in a Low-Neutron D–³He Tandem Mirror 100 次
Accelerating scientific discovery with Co-Scientist 94 次

量化金融方向

论文 被引
A Survey of Large Language Models 1550 次
Exploring Large Language Model‐Based Intelligent Agents: Definitions, Methods, and Prospects 38 次
Large Models for Time Series and Spatio-Temporal Data: A Survey and Outlook 35 次
The Scenario Model Intercomparison Project for CMIP7 (ScenarioMIP-CMIP7) 32 次
Investigating the replicability of the social and behavioural sciences 31 次

Hacker News 热门(科技 / 创业 / 投资风向)

热度 讨论 标题
708 赞 294 评 Exfiltrate your Weights
642 赞 455 评 What happened to the Snowden archive
619 赞 283 评 AX – Google’s Open Agentic Orchestrator
542 赞 421 评 Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
356 赞 189 评 What Sun got wrong
336 赞 75 评 [Grim Fandango Puzzle Document (1996) [pdf]
http://gameshelf.jmac.org/2008/11/13/GrimPuzzleDoc_small.pdf
- [331 赞 / 313 评论] ZuckOff is a free app that sees Meta glasses before they see you](https://www.wired.me/story/meta-smart-glasses-detector-app-zuckoff)
328 赞 153 评 Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
321 赞 94 评 Attention is all you have
310 赞 255 评 Grok 4.7
232 赞 49 评 Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM
205 赞 129 评 Fable 5 – Median thinking declined in August
185 赞 73 评 Heretic removes restrictions from language models
182 赞 155 评 M5 Ultra Mac Studio Review
171 赞 139 评 Raspberry Pi blocks changing RAM chips

本日报由服务器每日自动抓取生成于凌晨,原文链接均已附在上。