cs.CV Sep 10
Haiwen Diao, Jiahao Wang, Chenjing Ding +62
TL;DR: SenseNova-U1.5は、エンコーダーやVAEを使用しない8B-MoTネイティブ統一マルチモーダルモデルで、視覚コンテンツの理解、推論、生成を強化し、最大4Kの解像度で高い画像忠実度や複雑な構成を実現します。トレーニングコードはオープンソースとして提供され、視覚的計画と創造におけるマルチモーダル理解の可能性を示しています。
We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its vi...
cs.DC cs.AI cs.GT Sep 10
Boning Li, Longbo Huang
TL;DR: GPU-CFRは、カウンターファクチュアルレグレット最小化(CFR)を静的データフローとCUDAグラフリプレイを用いて80倍以上高速化する新しいコンパイラおよびランタイムであり、従来のGPU実装やCPU実装に対しても大幅な性能向上を実現しています。
Counterfactual regret minimization (CFR) is one of the few large numerical workloads that still runs faster on CPUs than on GPUs. Each iteration sweeps a game tree with up to billions of states in mil...
cs.LG cs.AI stat.ML Sep 10
Hongbo Chen, Li Charlie Xia
TL;DR: TL;DR: 本論文では、機械学習における分布シフトの一般的な定量化を提案し、既存の概念シフトの定義の限界を指摘した上で、$γ^{*}\!$-概念シフトを導入し、広範な損失関数やラベル空間に適用可能な一般的な誤差境界を導出した。また、分布シフトを定量化し誤差境界を推定するためのDataShiftsアルゴリズムを開発した。
Generalization under distribution shift remains a core challenge in modern machine learning, yet existing learning bound theory is limited to narrow, idealized settings and is non-estimable from sampl...
cs.LG cs.CL Sep 10
Atindra Jha, Margaret Li, Jure Leskovec +2
TL;DR: TL;DR: Mixture-of-Experts (MoE)モデルはデータの繰り返しに対して過剰適合しやすく、特にスパース性が高いほどその傾向が強まることが示されました。既存の正則化手法が過剰適合を軽減する可能性があるものの、全てのユニークなデータに匹敵する性能には至らず、今後の研究でメモリゼーションを妨げる方法の開発が期待されます。
As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but...
cs.AI Sep 10
William Zhou, Mayukha Siripuram, Xiao Yan +2
TL;DR: エッジデプロイ可能なビジョン・ランゲージモデル(VLM)が種の識別において有効かを評価した結果、全モデルがフィールド画像での識別精度が大幅に低下することが確認され、特に専門モデルのBioCLIPが他のVLMよりも優れた性能を示したが、画像の質の変化が識別能力の低下に寄与していることが示唆された。
Camera traps often run in the field on edge hardware with limited or no connectivity, making small, locally-deployable vision-language models (VLMs) -- not frontier-scale ones -- the practically relev...
stat.ML cs.AI cs.LG Sep 10
Masahiro Kato, Daiki Honma, Taka Kato
TL;DR: Generative Marketing Mix Modeling (GMMM)は、Generative Engine Optimization (GEO)とGenerative Engine Marketing (GEM)の因果効果を推定するフレームワークであり、生成された回答における企業名の認知度を考慮してビジネスへの影響を評価します。シミュレーションデータを用いて、提案手法の実証的な性能を検証しました。
Generative artificial intelligence changes how firms reach customers, but standard marketing data do not record how often users see and notice a firm's name in generated answers. We develop Generative...
cs.CL Sep 10
Daniel Henrik Nevermann, Claudius Gros
TL;DR: TL;DR: 本研究では、トランスフォーマーにおける距離一般化を探求し、位置エンコーディング(RoPEやALiBi)が距離解像度に与える影響や、トレーニング中に見たトークン間距離の多様性が性能に及ぼす効果を調査しました。結果として、距離転送学習のメカニズムを理解することが重要であることが明らかになりました。
Out-of-distribution length generalization, namely to extrapolate a task from short to longer context, has been studied intensively for transformers. Here we focus on distance generalization, which pro...
cs.AI Sep 10
Yakov Pyotr Shkolnikov
TL;DR: エージェントAIは、タスクの境界を越えて適応し続けるシステムへと進化しており、その制御問題を解決するために「人工的なID」を提案しています。このIDは、行動の継続、停止、変更を判断する内部ドライブを持ち、適応的なエージェンシーを実現する一方で、誤った整合性や意図しない行動が持続するリスクも伴います。
Agentic AI is moving from bounded task execution toward systems that retain consequential state, continue operating and adapt across task boundaries. That shift creates a control problem that current ...
cs.LG Sep 10
Nitesh V. Chawla, Paulo Benanti
TL;DR: AIはガバナンスの問題を引き起こすだけでなく、既存の制度の失敗を明らかにし、介入として機能する可能性がある。責任あるAIの実現には、制度の修復や倫理的判断が求められ、RISE AIは責任、包括性、安全性、エンパワーメントに基づく限界のある証拠に基づく主張を行うための枠組みを提供する。
Artificial Intelligence does more than create a governance problem. It can also reveal where institutions have already failed to provide responsiveness, belonging, care, and accountability. Once deplo...
cs.LG Sep 10
Akshaj Gupta, Hwi Joo Park, Andrea Guzman +5
TL;DR: TARTは、スライドやベンドなどの表現技法を捉え、正確な弦とフレットの割り当てを行い、ノイズの多い音声でも機能するギター音声からタブラチュアへの自動転写を実現するモジュラー型ツールであり、評価結果では従来のベースラインを上回る性能を示しました。
Automatic Music Transcription (AMT) for guitar remains limited by three challenges: existing systems often fail to capture expressive techniques such as slides, bends, and percussive hits; they often ...
cs.AI cs.CL cs.CV Sep 10
Yunfei Ge, Anbang Liu, Qineng Wang +9
TL;DR: MindTopoは、連続変形に対して不変なトポロジー関係を評価するベンチマークであり、空間的推論における基礎的な理解を測定するために、連続性、分離、順序、囲い、結び目の5つの特性を用いています。14のMLLMを評価した結果、推論では人間のパフォーマンスには及ばないものの、計画よりも良好な結果を示しました。
Spatial reasoning depends not only on metric properties such as distance, angle, and shape, but also on topological relations that remain invariant under continuous deformation. Cognitive science iden...
cs.CV cs.HC Sep 10
Weitong Cai, Hang Zhang, Yukai Huang +6
TL;DR: 長動画理解のための新しいフレームワーク「Caption-once, Frames-on-Demand(CFD)」は、エッジデバイス上での計算資源と帯域幅の制約を考慮し、視覚的ニーズに応じたフレーム取得を行うことで、効率的な情報処理を実現します。このアプローチにより、長期的な時間構造を保持しつつ、必要な視覚情報のみを効果的に取得することが可能になります。
Long-video understanding on edge devices must reason over hours of content under tight compute and bandwidth budgets. Subsampling visual tokens loses temporal structure, while text-only video memories...