Library

Subject
Tags

24,106 matches · Vision

#
001Bowel Obstruction Detection and Localization on Abdominal CT with Deep LearningarXivPaperMoritz Vandenhirtz, Andrea Agostini et al.7 days ago
002CARDIAG: A Dense Segment Classification Benchmark of Deep Learning Architectures for Coronary AngiographyarXivPaperDominik Bernard Lau, Hubert Malinowski et al.7 days ago
003Class-Balanced Softmax: A Bayes Theory-Based Method for Long-Tailed RecognitionarXivPaperYi-Hang Zhu, Rajeev Raman et al.7 days ago
004Correlation-Aware and Gaussianity-Preserving Robust Latent Angular Watermarking for Diffusion ModelsarXivPaperYebin Zheng, Haonan An et al.7 days ago
005Deep Convolutional Large-Margin $\ell_p$-SVDD for Visual Anomaly DetectionarXivPaperAlireza Dastmalchi Saei, Shervin Rahimzadeh Arashloo7 days ago
006Diffusion Models in Medical Image Inpainting: Challenges, Solution Taxonomy, and Future DirectionsarXivPaperErik C. Rye, Robert Beverly7 days ago
007EVL-MCoT: Enhanced Vision-Language Multi-CoT for Harmful Meme DetectionarXivPaperHao Yang, Jin Wang et al.7 days ago
008Farmland Extent and Visible Boundary Mapping from 1 m NAIP Imagery Using Residual U-Net and Text-Prompted SAM 3 RefinementarXivPaperSreejeet Maity, Feng Zhu et al.7 days ago
009Filling Before Advancing: Capability-Gap-Driven Post-Training for Scenario-Specialized Remote Sensing MLLMsarXivPaperYuheng Zong, Minghua Wang et al.7 days ago
010ISPCloak: Weaponizing ISP for Optimization-Free Physical Camouflage against Deepfake DetectorsarXivPaperLiangqin Ren, Zeyan Liu et al.7 days ago
011Rethinking Multi-Branch and Cross-Backbone Fusion for Vehicle Re-Identification in the Foundation-Model EraarXivPaperYu Wang, Hongyu Yang7 days ago
012SM4RT: Learning Structured Motion Geometry for 4D ReconstructionarXivPaperShing Ho J. Lin, Wenzhao Zheng et al.7 days ago
013Scaling Native Multimodal Pre-Training From ScratcharXivPaperHaoyuan Wu, Aoqi Wu et al.7 days ago
014SceneActBench: Can Agents Act on the 3D Scenes They See?arXivPaperYifei Zhao, Xiangxin Zhou et al.7 days ago
015SiPhy: Single-Image Physical Property ReasoningarXivPaperH. Lê, Joon-Byum Kwon et al.7 days ago
016Small Vision-Language Models Know When They Are Wrong But Cannot Say So: A Two-Model Study of Stated versus Internal Confidence Under Realistic Image DegradationarXivPaperAsif Ferdous7 days ago
017TRaM-VSR: Importance-Aware Token Routing and Merging for One-Step Diffusion Video Super-ResolutionarXivPaperSicheng Gao, Zhuyun Zhou et al.7 days ago
018TextSLIP: Text Self-Supervised CLIP for Medical Report GenerationarXivPaperShashank Rao Marpally, Allan Wang et al.7 days ago
019Time-Reversed Imaging: A Multimodal Benchmark and Framework for Reconstructing Past Human-Environment InteractionsarXivPaperJorge Bacca, Kebin Contreras et al.7 days ago
020Unboxing Diffusion Models for the Arts: Interactive Model Bending and Practice-Based ExplainabilityarXivPaperAhmed M. Abuzuraiq, Philippe Pasquier7 days ago
021Visual Saliency Steering Distillation for Multimodal Chain-of-Thought ReasoningarXivPaperHao Yang, Jin Wang et al.7 days ago
022Zero-Shot Mission-Level Evaluation for Aerial MLLM AgentsarXivPaperSuman Navaratnarajah, Taehyoung Kim et al.7 days ago
023dRAE: Representation Autoencoder with Hyper-Spherical CodesarXivPaperTianren Ma, Lin Long et al.7 days ago
0243D-Aware VLMs with Implicit and Explicit GeometriesarXivPaperWenhao Li, Xueying Jiang et al.Jul 23
025Adaptive Identity Anchoring: Closed-Loop Keyframe Placement for Synthetic Paired Supervision in Video Face SwappingarXivPaperLogan RobbinsJul 23
026CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQAarXivPaperHanseok Oh, Parishad BehnamGhader et al.Jul 23
027Counterfactual Explainability Framework With CycleGAN And Counterfactual-Classifier Alignnment Score for Retinal Disease ClassificationarXivPaperKritanu Chattopadhyay, Sayanjit Singha Roy et al.Jul 23
028DART: A Degradation-Aware Recurrent Transformer for Archival Film RestorationarXivPaperMikołaj Jastrzębski, Wojciech Kozlowski et al.Jul 23
029DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic SegmentationarXivPaperSung-Hoon Yoon, Hoyong Kwon et al.Jul 23
030ElasticTTT: Prior-Preserving Test-Time Tuning for Video EditingarXivPaperYueyi Liu, Chi Zhang et al.Jul 23
031EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent SpecializationarXivPaperLihuang Fang, Yuchen Zou et al.Jul 23
032GS-Agent: Creating 4D Physical Worlds With Generative SimulationarXivPaperHongxin Zhang, Chunru Lin et al.Jul 23
033GraphVid: Interactive Graph-Controllable Video GenerationarXivPaperVedant r Shah, Onkar Susladkar et al.Jul 23
034HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous DrivingarXivPaperQuanfu Yu, Xian Wu et al.Jul 23
035Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla FormsarXivPaperXiao Zhu, S. IyengarJul 23
036KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion TransformersarXivPaperYann Bouquet, Alireza Khodamoradi et al.Jul 23
037M$^3$-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging DataarXivPaperFrancesca Pia Panaccione, Carlo Sgaravatti et al.Jul 23
038PC-Edit: Prompt-Contrastive Region Discovery and Region-Guided EditingarXivPaperJian Zhang, Zhijun ZhangJul 23
039Physiological Signals as a Forensic Modality for Talking-Face Deepfake DetectionarXivPaperPablo Santiago Potes Velasco, María del Mar García Matabanchoy et al.Jul 23
040Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and ImagesarXivPaperIdris Karel Seunda Ekwe, Patrick Tenga Shako et al.Jul 23
041SCALE: Self-Supervised Constraint-Aware Layout GEneration for Local P&R DRV Fixing at Advanced NodesarXivPaperKhai Nguyen, Yang Ni et al.Jul 23
042Safety-oriented sidewalk and road segmentation for smartphone-based assistive navigationarXivPaperHakan Calim, Anamaria Dumitrescu et al.Jul 23
043Self-Poisoning in Adaptive Out-of-Distribution Detection: A Sharp-Threshold Theory and Certified Label-Free CalibrationarXivPaperRonak BhalgamiJul 23
044Sparse Concept Channels in Frozen 3D CT Vision EncodersarXivPaperF. Nooralahzadeh, L. Bogensperger et al.Jul 23
045Synthetic data generation framework for quality control automation in gravure printingarXivPaperKorota Arsène Coulibaly, Mohamed Hamlich et al.Jul 23
046Unlearning Under Imbalance: Benchmarking Fairness in Multimodal LLM UnlearningarXivPaperL. Orsingher, Thomas De Min et al.Jul 23
047Visual Contrastive Self-DistillationarXivPaperYijun Liang, Yunjie Tian et al.Jul 23
048When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time ModerationarXivPaperDongbin NaJul 23
049A Systematic Benchmark of Intensity Normalisation Methods for 3D Knee MRI Segmentation and Cross-Domain GeneralisabilityarXivPaperOliver Mills, Philip G. Conaghan et al.Jul 22
050Analytic Distribution of Classifier-Free Guidance for Schedule DesignarXivPaperEnze Jiang, Zheng MaJul 22
051Can an AI System Be Creative? A Critical Perspective from Art and EngineeringarXivPaperIvan Magrin-ChagnolleauJul 22
052DS@GT ARC at ImageCLEFmed GANs 2026: Geometric Filtering for Privacy-Preserving CT Slice GenerationarXivPaperEric Regina, Richard Arnaud et al.Jul 22
053Detecting Neural Network Failures through Spectral Analysis of Internal ActivationsarXivPaperAndy Song, Bolong Han et al.Jul 22
054ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language ModelsarXivPaperKaran Goyal, Afreen Hossain et al.Jul 22
055G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object DetectionarXivPaperYechan Kim, Jonghyun Park et al.Jul 22
056HeadCast: Casting Attention Heads for Efficient Autoregressive Video GenerationarXivPaperJinliang Shen, Li Su et al.Jul 22
057Masked Topology Modeling for Self-Supervised Learning on Parametric CADarXivPaperHeinrich Jiang, J. JangJul 22
058Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial VideosarXivPaperPenglei Sun et al.Jul 22
059OSVE: One Step Video Editing with One Step Diffusion ModelsarXivPaperHabin Lim, Gyeong-Moon ParkJul 22
060PRISM-DR: Per-lesion Retinal Inference with Specialist Models for Diabetic RetinopathyarXivPaperZ. Özeren, Tansel UyarJul 22

Showing 60 of 24,106 documents · scroll for more