Chemical Dice Integrator Unifies Molecular Information for Accessible Drug Discovery
Research Summary: Chemical Dice Integrator learns a unified molecular representation from six complementary data modalities and distills it into a SMILES-based model for scalable, robust drug-discovery prediction.
Researcher Spotlight
First authors name: Suvendu Kumar, Saveena Solanki, Mudit Gupta, Sonam Chauhan, Sanjay Kumar Mohanty.
Suvendu Kumar, Saveena Solanki, and Sonam Chauhana are Phd at IIIT-Delhi, are working at the intersection of deep learning, molecular representation learning, genomics, and AI-driven drug discovery. Mudit Gupta, a former BTech Student of IIIT-Delhi currently working as Software Development Engineer at Microsoft. Dr. Sanjay Kumar Mohanty is a Postdoctoral Researcher at the Indian Institute of Science (IISc), working in Neuroscience.
LinkedIn:
Suvendu Kumar https://in.linkedin.com/in/suvendu-kumar-50735717a
Saveena Solanki https://www.linkedin.com/in/saveenasolanki/
Mudit Gupta https://www.linkedin.com/in/mudit-gupta-375a18212/
Sonam Chauhan https://www.linkedin.com/in/sonam-chauhan-a56138170/
Sanjay Kumar Mohanty https://www.linkedin.com/in/sanjay-741mohanty/
Lab PI name: Dr. Gaurav Ahuja
University or institute: Indraprastha Institute of Information Technology, Delhi
Dr. Gaurav Ahuja LinkedIn: https://in.linkedin.com/in/gaurav-ahuja-45a80528
Lab Website https://www.ahuja-lab.in/
What was the core problem you aimed to solve with this research?
Most molecular artificial intelligence models see a chemical compound through only one lens, such as its structure, physicochemical properties, or biological activity. Each lens captures useful but incomplete information, while combining many data types directly is computationally expensive and often fails when one modality is unavailable. We aimed to create a unified molecular representation that retains complementary information, remains robust to missing source features, and can be generated from a simple SMILES string.

How did you go about solving this problem?
We developed Chemical Dice Integrator, or CDI, around six complementary molecular modalities: Mordred physicochemical descriptors, GROVER graph features, ImageMol visual features, Signaturizer bioactivity features, MOPAC quantum properties, and ChemBERTa chemical-language features. A hierarchical autoencoder first learns shared information across these views and fuses them into an 8,192-dimensional CDI-Basic representation. We then trained a Mamba state-space model to distill this rich representation directly from canonical SMILES, producing CDI-Generalized. We evaluated the embeddings across 23 classification datasets comprising 171 tasks, 10 regression datasets, scaffold-based out-of-distribution settings, limited-data experiments, modality ablations, retrieval analyses, and external chemical libraries.
“CDI turns complementary molecular knowledge into an accessible representation that can support more reliable and scalable chemical discovery.” – Saveena Solanki (Corresponding Author)
How would you explain your research outcomes (Key findings) to the non-scientific community?
A molecule can be described in several different ways, much like a person can be described by appearance, behaviour, history, and skills. CDI learns from several such descriptions together and compresses them into one reusable molecular fingerprint. Its distilled version needs only the molecule’s SMILES text, yet retains much of the broader information learned from the six original views. Across diverse prediction tasks, this unified representation was robust and broadly useful, while CDI-Generalized generated valid embeddings for every valid SMILES tested in the evaluated external libraries.
“CDI transforms diverse molecular perspectives into a unified representation, enabling robust predictions and accelerating practical chemical discovery across data-limited settings.” – Prof. Gaurav Ahuja (Corresponding Author)
What are the potential implications of your findings for the field and society?
CDI provides a practical bridge between information-rich molecular modeling and real-world deployment. Researchers can use a single SMILES string to obtain a reusable embedding for toxicity, bioactivity, physicochemical-property prediction, similarity search, and compound prioritization. This could help teams screen large chemical libraries more efficiently and focus laboratory effort on stronger candidates. The method is intended to support, not replace, experimental validation; its societal value will depend on careful prospective testing, transparent reporting, and responsible use in drug discovery.
What was the exciting moment during your research?
The most exciting moment was seeing CDI-Generalized reproduce the information-rich embedding from SMILES alone and maintain complete embedding coverage across chemically diverse external libraries. That result showed that the work had moved beyond a complex multimodal training system into something researchers could realistically deploy, including when upstream quantum or pretrained encoders fail for individual molecules.
Paper reference and project links: Kumar, S., Solanki, S., Gupta, M. et al. Scalable molecular representations enabled by multimodal fusion and sequence distillation. Nat Commun (2026). https://doi.org/10.1038/s41467-026-77700-z .


