Skip to content

๐Ÿ”— Causal Inference

๐Ÿ“น ICCV2025 ยท 2 paper notes

๐Ÿ“Œ Same area in other venues: ๐Ÿ“ท CVPR2026 (4) ยท ๐Ÿ”ฌ ICLR2026 (64) ยท ๐Ÿ’ฌ ACL2026 (7) ยท ๐Ÿงช ICML2026 (19) ยท ๐Ÿค– AAAI2026 (7) ยท ๐Ÿง  NeurIPS2025 (20)

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets

This paper proposes a block-based diffusion method leveraging LLMs and diffusion models to automatically generate high-quality counterfactual image-text pair datasets, accompanied by a set-aware loss function. Without manual annotation, the approach significantly improves CLIP's compositional reasoning ability, surpassing state-of-the-art methods on ARO/VL-Checklist and other benchmarks with substantially less data.

Social Debiasing for Fair Multi-modal LLMs

This paper constructs CMSC, a large-scale counterfactual dataset spanning 18 social concepts, and proposes the Anti-Stereotype Debiasing (ASD) strategyโ€”comprising bias-aware data resampling and a Social Fairness Lossโ€”that effectively reduces social bias across four MLLM architectures with negligible degradation of general multimodal capability.