Weakly Supervised Concept Learning for Object-centric Visual Reasoning
Researchers from arXiv have introduced a novel weakly supervised concept learning framework designed to enhance object-centric visual reasoning within neurosymbolic systems. This approach addresses the high cost of labeled data in two-stage architectures by decoupling deep neural network perception from symbolic reasoning. The method combines a slot-based architecture for object-centricity with a Variational Autoencoder (VAE) for self-supervision, enabling human-interpretable grounding of output symbols. These predictions are subsequently translated into symbolic background knowledge compatible with reasoning frameworks like Inductive Logic Programming, Decision Trees, and Bayesian Networks. Ext empirical evaluations on both synthetic and real-world datasets demonstrate that the system can discover complex, abstract rules while reducing supervision requirements to merely 1% of labels. Notably, the model exhibits robustness against substantial domain shifts and outperforms state-of-the-art foundation model baselines in domain generalization tasks under minimal supervision conditions. This advancement promises to improve the efficiency and interpretability of AI systems requiring logical induction from raw sensor inputs.
Wire timeline
Weakly Supervised Concept Learning for Object-centric Visual Reasoning
Researchers from arXiv have introduced a novel weakly supervised concept learning framework designed to enhance object-centric visual reasoning within neurosymbolic systems. This approach addresses the high cost of labeled data in two-stage architectures by decoupling deep neural network perception from symbolic reasoning. The method combines a slot-based architecture for object-centricity with a Variational Autoencoder (VAE) for self-supervision, enabling human-interpretable grounding of output symbols. These predictions are subsequently translated into symbolic background knowledge compatible with reasoning frameworks like Inductive Logic Programming, Decision Trees, and Bayesian Networks. Ext empirical evaluations on both synthetic and real-world datasets demonstrate that the system can discover complex, abstract rules while reducing supervision requirements to merely 1% of labels. Notably, the model exhibits robustness against substantial domain shifts and outperforms state-of-the-art foundation model baselines in domain generalization tasks under minimal supervision conditions. This advancement promises to improve the efficiency and interpretability of AI systems requiring logical induction from raw sensor inputs.
cs.AI updates on arXiv.org