2026
Towards Robustness against Typographic Attack with Training-free Concept Localization
Bohan Liu, Wenqian Ye, Guangzhi Xiong, Zhenghao He, Sanchit Sinha, Aidong Zhang
ECCV 2026 paper ↗

Text painted on an image shouldn't decide what the model sees. We localize, training-free, the small set of CLIP attention heads that read injected words — then reweight or ablate them. +22.8 points object accuracy on RTA-100 (ViT-H/14) with ~zero test-time cost, beating supervised and training-free defenses across five backbones.
MM-SPUBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
Wenqian Ye, B Liu, G Zheng, D Wang, X Cao, Y Ma, B Lai, JM Rehg, A Zhang
A benchmark for spurious biases in multimodal LLMs: paired image sets where the correct answer is independent of spurious cues (backgrounds, text overlays, co-occurrences). Reveals that strong MLLMs lean heavily on shortcuts, motivating bias-aware evaluation and mitigation.
Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning
S Sinha, G Xiong, B Liu, Z He, A Zhang
arXiv Jun 2026 paper ↗

Fine-tuning only a few attention maps — not weights — of multimodal LLMs improves chain-of-thought reasoning. Attention-guided updates steer where the model looks while keeping its knowledge intact, lifting ChartQA and CoT benchmarks at a fraction of full fine-tuning cost.
Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
G Xiong, Z He, B Liu, S Sinha, A Zhang
RAG models often ignore retrieved evidence when generating. We use sparse autoencoders to decompose generation behavior, identify features responsible for faithfulness, and steer them — improving attribution and grounding without retraining the base model.
SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias
W Ye, D Wang, G Zheng, B Liu, A Zhang
Prompt-level defense against multimodal bias: SAGE estimates how 'spuriousness-aware' a prompt is and explores guided prompt variations to find framings where the model relies on true visual content rather than dataset shortcuts.
Reasoning Beyond Chain-of-Thought: A Latent Computational Mode in Large Language Models
Z He, G Xiong, B Liu, S Sinha, A Zhang

Chain-of-thought is not the only way LLMs reason. We uncover a latent computational mode — reasoning that completes inside hidden states before any tokens are emitted — and characterize when latent computation beats explicit CoT.
2025
UrbanIR: Large-Scale Urban Scene Inverse Rendering from a Single Video
CH Lin, B Liu, YT Chen, KS Chen, D Forsyth, JB Huang, A Bhattad, et al.
Inverse rendering of large-scale urban scenes from a single casually-captured video: decomposing the city into albedo, geometry, and transient lighting, enabling photorealistic relighting and editing far beyond prior single-image methods.
2022
Online Classifier of AMICA Model to Evaluate State Anxiety While Standing in Virtual Reality
G Liao, S Wang, Z Wei, B Liu, R Okubo, ME Hernandez
An online classifier built on AMICA EEG features that detects state anxiety while a person stands in virtual reality — real-time monitoring with applications in fall-prevention and VR therapy.