Daily Paper Cast

Daily Paper Cast

byJingwen Liang, Gengyu Wang

ScienceTechnology

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.com Creator: Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/ Gengyu Wang, LLM ML, http://wanggengyu.com Listen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXL Apple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236 Cover Image by Kawen Kuang https://kawen.art

Episodes(40 episodes)

Episode 2067
Apple-$Ο€$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
πŸ€— Upvotes: 39 | cs.CV Authors: Runmao Yao, Kairui Hu, Yukang Cao, Ruisi Wang, Shulin Tian, Ziang Cao, Weichen Fan, Ziqi Huang, Yuhao Dong, Hao Li, Zhaoxi Chen, Zhongang Cai, Lei Yang, Ziwei Liu Title: Apple-$Ο€$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Arxiv: http://arxiv.org/abs/2607.16401v1 Abstract: Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through a fa...
Published: Jul 22, 2026Duration: 22m 3s
Episode 2066
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
πŸ€— Upvotes: 119 | cs.SE, cs.AI Authors: Yijia Fan, Zonglin Di, Zimo Wen, Yifan Yang, Mingxi Cheng, Qi Dai, Bei Liu, Kai Qiu, Yue Dong, Ji Li, Chong Luo Title: RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources Arxiv: http://arxiv.org/abs/2606.29538v4 Abstract: Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing skill libraries are mostly hand-written, text-centric, or derived from agent traces, leaving tutorial videos and other multimodal human resources largely underused. We pre...
Published: Jul 21, 2026Duration: 19m 13s
Episode 2065
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM
πŸ€— Upvotes: 115 | cs.CL, cs.AI Authors: Mikhail Komarov, Ivan Bondarenko, Stanislav Shtuka, Oleg Sedukhin, Roman Shuvalov, Yana Dementyeva, Matvey Solovyov, Nikolay O. Nikitin Title: RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM Arxiv: http://arxiv.org/abs/2607.11683v1 Abstract: Graph retrieval-augmented generation (GraphRAG) enhances large language models with structured knowledge, yet existing systems construct knowledge graphs in a single extraction pass, producing noisy entities and brittle retrieval. RAGU, an open-source modular GraphRAG engine, addresses this by separating extraction from consolidation: entities and relations pass through two...
Published: Jul 21, 2026Duration: 19m 44s
Episode 2064
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
πŸ€— Upvotes: 57 | cs.RO, cs.CV Authors: Xiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li, Peiyan Li, Yingyan Li, Futeng Liu, Wanli Peng, Optimus Qin, Yifei Su, Nan Sun, Qiao Sun, Runze Suo, Heyun Wang, Yunhong Wang, Rujie Wu, Caoyu Xia, Lina Zhang, Jack Zhao, Guoliang Chen, Wenlong Chen, Xinze He, Bin Li, Qing Li, Zhuorong Li, Heng Qu, Wenxuan Song, Diyun Xiang, Yifan Xie, Peiran Xu, Hangjun Ye, Wen Ye, Han Zhao, Quanyun Zhou Title: Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Arxiv: http://arxiv.org...
Published: Jul 21, 2026Duration: 20m 54s
Episode 2063
Loop the Loopies!
πŸ€— Upvotes: 56 | cs.CL, cs.AI Authors: Zitian Gao, Yilong Chen, Yihao Xiao, Xinyu Yang, Ran Tao, Joey Zhou, Bryan Dai Title: Loop the Loopies! Arxiv: http://arxiv.org/abs/2607.16051v2 Abstract: We present the Loopie series, consisting of two Mixture-of-Experts (MoE) models: a 20B-parameter model with 2B active parameters and a 6B-parameter model with 0.6B active parameters. Looped Transformers have long faced a challenge: given an N times increase in pre-training compute, increasing the parameter count by a factor of N usually outperforms looping a model N tim...
Published: Jul 21, 2026Duration: 19m 35s
Episode 2062
xHC: Expanded Hyper-Connections
πŸ€— Upvotes: 47 | cs.LG, cs.CL Authors: Xiangdong Zhang, Xiaohan Qin, Sunan Zou, Tuo Dai, Xiaoming Shi, Huaijin Wu, Yebin Yang, Zhuo Xia, Shaofeng Zhang, Lin Yao, Yuliang Liu, Yu Cheng, Junchi Yan Title: xHC: Expanded Hyper-Connections Arxiv: http://arxiv.org/abs/2607.14530v1 Abstract: Hyper-Connections (HC) expand the residual stream of Transformers into $N$ parallel streams, providing a form of memory scaling beyond model width and depth. Manifold-Constrained HC (mHC) stabilizes this formulation at scale. The large gains from $N{=}1$ to $N{=}4$ suggest residual-stream expansion as a promising sca...
Published: Jul 21, 2026Duration: 20m 38s
Episode 2061
Cura 1T: Specialized Model for Agentic Healthcare
πŸ€— Upvotes: 43 | cs.AI Authors: actAVA AI, :, Haolin Chen, Leon Qi, Steve Brown, Deon Metelski, Tao Xia, Joonyul Lee, Qixuan Wang, Kevin Riley, Frank Wang, Weiran Yao Title: Cura 1T: Specialized Model for Agentic Healthcare Arxiv: http://arxiv.org/abs/2607.15314v1 Abstract: Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover these use cases together remain limited. A healthcare model must handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use. These capabilities fail in dif...
Published: Jul 21, 2026Duration: 20m 49s
Episode 2060
On-Policy Delta Distillation
πŸ€— Upvotes: 28 | cs.LG, cs.CL Authors: Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han Title: On-Policy Delta Distillation Arxiv: http://arxiv.org/abs/2607.15161v1 Abstract: On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constraints imposed by reward models by providing token-level supervision from a teacher model. Although on-policy distillation has been studied and applied across various settings, its fundamental design remains underexplored. In this paper, we introduce a new distillation reward, termed the delta signal, instead of directly imitating the teacher's output dis...
Published: Jul 21, 2026Duration: 19m 41s
Episode 2059
RecGPT-V3 Technical Report
πŸ€— Upvotes: 26 | cs.IR Authors: Bowen Zheng, Chao Yi, Dian Chen, Gaoyang Guo, Han Zhu, Jiakai Tang, Jian Wu, Mao Zhang, Wen Chen, Yifan Lu, Yujie Luo, Yuning Jiang, Zhujin Gao, Bo Zheng, Dixuan Wang, Hao Fang, Jiancai Liu, Jing Yu, Ke Chen, Kewei Zhu, Mingke Xu, Wenjun Yang, Xunke Xi, Zile Zhou Title: RecGPT-V3 Technical Report Arxiv: http://arxiv.org/abs/2607.15591v1 Abstract: Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in historical behavior toward reasoning about the intent that drives it. RecGPT-V1 pio...
Published: Jul 21, 2026Duration: 22m 0s
Episode 2058
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
πŸ€— Upvotes: 109 | cs.CV Authors: Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, Haoning Wu, Zhiqiu Zhang, Yuandong Yang, Changlian Ma, Qingyu Zhang, Yansong Shi, Xinyu Chen, Haoran Chen, Zizheng Huang, Jun Zhang, Kun Ouyang, Lin Sui, Ziang Yan, Yicheng Xu, Chenting Wang, Yinan He, Hongjie Zhang, Yi Wang, Yu Qiao, Yali Wang, Ziwei Liu, Kai Chen, Limin Wang Title: VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Arxiv: http://arxiv.org/abs/2607.14935v1 Abstract: Recent advances in video understanding have spanned motion, long video, and...
Published: Jul 18, 2026Duration: 24m 41s
Episode 2057
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
πŸ€— Upvotes: 71 | cs.CL Authors: Jinyang Wu, Shuo Yang, Zhengxi Lu, Fan Zhang, Yuhao Shen, Lang Feng, Haoran Luo, Zheng Lian, Shuai Zhang, Zhengqi Wen, Jianhua Tao Title: SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Arxiv: http://arxiv.org/abs/2607.14777v1 Abstract: Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimization paradigm, but its sparse trajectory-level rewards offer limited guidance on intermediate decisions, leaving a supervision gap between epi...
Published: Jul 18, 2026Duration: 19m 16s
Episode 2056
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
πŸ€— Upvotes: 83 | cs.LG, cs.DC Authors: Changhai Zhou, Kieran Liu, Yuhua Zhou, Qian Qiao, Jun Gao, Harry Zhang, Irvine Lu, Nolan Ho, Lucian Li, Andrew Lei, Cleon Cheng, Steven Chiang, Yihang Zeng, Di Zhang, Rio Yang, Kaijie Chen, Andrew Chen, Pony Ma, Weizhong Zhang, Cheng Jin Title: LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Arxiv: http://arxiv.org/abs/2607.14952v1 Abstract: A growing gap separates inference context lengths from RL post-training: inference systems are approaching million-token contexts, while post-training workloads often remain at 256K t...
Published: Jul 18, 2026Duration: 19m 36s
Episode 2055
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
πŸ€— Upvotes: 49 | cs.AI, cs.IR Authors: Yuyao Zhang, Junjie Gao, Zhengxian Wu, Jiaming Fan, Jin Zhang, Shihan Ma, Yao Yao, Weiran Qi, Chuyan Jin, Guiyu Ma, Xingzhong Xu, Kai Yang, Ji-Rong Wen, Zhicheng Dou Title: SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration Arxiv: http://arxiv.org/abs/2607.15257v1 Abstract: Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seeking agents. However, as interaction histories grow, agents increasingly struggle to track task progress. When search attempts fail to yield useful evidence, current sin...
Published: Jul 18, 2026Duration: 15m 57s
Episode 2054
BadWAM: When World-Action Models Dream Right but Act Wrong
πŸ€— Upvotes: 36 | cs.LG, cs.RO Authors: Qi Li, Xingyi Yang, Xinchao Wang Title: BadWAM: When World-Action Models Dream Right but Act Wrong Arxiv: http://arxiv.org/abs/2607.15207v1 Abstract: World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we sho...
Published: Jul 18, 2026Duration: 21m 3s
Episode 2053
KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation
πŸ€— Upvotes: 30 | cs.CV Authors: Yuqi Tang, Tengfei Liu, Yizheng Lai, Yuran Wang, Yang Shi, Wanshun Su, Zhuoran Zhang, Qixun Wang, Xiaohan Zhang, Xinlei Yu, Xuehai Bai, Xuanyu Zhu, Bohan Zeng, Bozhou Li, Shujie Li, Yifan Dai, Yujie Wei, Shixuan Liu, Haotian Wang, Jialu Chen, Yuanxing Zhang Title: KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation Arxiv: http://arxiv.org/abs/2607.14202v1 Abstract: Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it rem...
Published: Jul 18, 2026Duration: 22m 20s
Episode 2052
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
πŸ€— Upvotes: 29 | cs.CV, cs.SD Authors: Xiaohan Zhang, Yuqing Wen, Junlin Chen, Yuqi Tang, Yiting He, Lizhuo Shao, Weiming Zhu, Tengfei Liu, Yang Shi, Jialu Chen, Yuanxing Zhang, Huaxiong Li Title: MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation Arxiv: http://arxiv.org/abs/2607.14189v1 Abstract: Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated audio-video alignment, leaving the emerging MR2AV setting largely unexplored. Compared with these set...
Published: Jul 18, 2026Duration: 20m 51s
Episode 2051
Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes
πŸ€— Upvotes: 23 | cs.LG Authors: Minh-Quan Le, Armand Comas, Alexandros Lattas, Stylianos Moschoglou, Pedro VΓ©lez, Amit Raj, Aaron Germuth, Thabo Beeler, Dimitris Samaras, Di Qiu Title: Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes Arxiv: http://arxiv.org/abs/2607.13188v1 Abstract: Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws $\textit{together}$, each modality reshapes the other. In this paper, we bring this coupled loop to artificial systems. Masked Diffusion Models (MDMs) are ideally suited to this task, yet...
Published: Jul 18, 2026Duration: 21m 40s
Episode 2050
From Pixels to States: Rethinking Interactive World Models as Game Engines
πŸ€— Upvotes: 23 | cs.CV Authors: Zhen Li, Zian Meng, Shuwei Shi, Mingliang Zhai, Jiaming Tan, Chuanhao Li, Kaipeng Zhang Title: From Pixels to States: Rethinking Interactive World Models as Game Engines Arxiv: http://arxiv.org/abs/2607.14076v1 Abstract: Building interactive worlds that respond coherently to player actions has long been a shared goal of computer graphics, games, and artificial intelligence. Recent video generative models provide a data-driven route toward this goal by predicting future observations conditioned on user actions, and are increasingly regarded as potential next-generation game engines. Rea...
Published: Jul 18, 2026Duration: 17m 56s
Episode 2049
UniVR: Thinking in Visual Space for Unified Visual Reasoning
πŸ€— Upvotes: 22 | cs.CV Authors: Zhongwei Ren, Yunchao Wei, Yao Zhao, Weibo Gong, Xiao Liu, Anran Wang, Xiangtai Li, Xiaojie Jin Title: UniVR: Thinking in Visual Space for Unified Visual Reasoning Arxiv: http://arxiv.org/abs/2607.12800v1 Abstract: Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations. At its core, UniVR features VR-GRPO, a reinforcement learning paradigm with complementary global and ste...
Published: Jul 18, 2026Duration: 18m 17s
Episode 2048
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
πŸ€— Upvotes: 172 | cs.AI, cs.SE Authors: Ruhan Wang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Yue Yu, Junyao Yang, Kishan Panaganti, Haitao Mi, Dongruo Zhou, Leoweiliang Title: Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Arxiv: http://arxiv.org/abs/2607.13285v1 Abstract: The capability of a modern AI agent depends not only on its foundation model but also on its harness, which constructs prompts, manages state, invokes tools, and coordinates execution. As models, APIs, environments, and requirements evolve, the harness must be continually modified. Before such a c...
Published: Jul 17, 2026Duration: 19m 12s