Discover how to build a comprehensive multimodal augmentation and adversarial robustness workflow using AugLy for images, text, audio, and PyTorch…
Tag: AI
At Meta Connect, the company’s smart glasses were everywhere
The company behind Facebook and Instagram wants to keep consumers connected to the digital world via its ever-growing line of smart glasses.
ADATEX4D: adaptive texture capacity allocation for 4D gaussian splatting
arXiv:2609.29963v1 Announce Type: cross Abstract: Textured Gaussians improve local appearance capacity, but assigning the same texture resolution to every…
Beyond Average Safety: Chance-Constrained LLM Fine-tuning
arXiv:2609.29960v1 Announce Type: cross Abstract: Fine-tuning large language models on new objectives can improve helpfulness, instruction following, or…
From Interests to Semantic IDs: Retrieval-Grounded Credit Assignment for Generative Recommendation
arXiv:2609.29983v1 Announce Type: cross Abstract: Semantic IDs (SIDs) encode each catalog item as a short token sequence, enabling generative recommenders…
World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal
arXiv:2609.29964v1 Announce Type: cross Abstract: General-purpose vision-language models (VLMs) bring broad knowledge and spatial reasoning to robot…
Learning Better Reasoning for Generative Recommendation with Semantic IDs
arXiv:2609.29973v1 Announce Type: cross Abstract: Generative recommendation reformulates item retrieval as sequence generation, allowing a unified model…
Tracking States or Tracking Cosets? An Algebraic Account of Learned State Tracking
arXiv:2609.29951v1 Announce Type: cross Abstract: State tracking requires composing a sequence of updates, but accuracy alone does not reveal what a model…
Mind What Matters for Reasoning: Aligning Cross-Modal Attention via Selective Probability Mass Concentration
arXiv:2609.29940v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) achieve strong performance on visual reasoning tasks, yet…
An Empirical Study of VLM Pipelines for Long-Document QA
arXiv:2609.29933v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly used for long-document processing, where the inputs…
