Singh Sisters: Master’s, PhD & Postdoc Guidance | Study Abroad

Chemical Dice Integrator Unifies Molecular Data for AI-Driven Drug Discovery

Chemical Dice Integrator Unifies Molecular Information for Accessible Drug Discovery

Research Summary: Chemical Dice Integrator learns a unified molecular representation from six complementary data modalities and distills it into a SMILES-based model for scalable, robust drug-discovery prediction.

Researcher Spotlight

First authors name: Suvendu Kumar, Saveena Solanki, Mudit Gupta, Sonam Chauhan, Sanjay Kumar Mohanty.

Suvendu Kumar, Saveena Solanki, and Sonam Chauhana are Phd at IIIT-Delhi, are working at the intersection of deep learning, molecular representation learning, genomics, and AI-driven drug discovery. Mudit Gupta, a former BTech Student of IIIT-Delhi currently working as Software Development Engineer at Microsoft. Dr. Sanjay Kumar Mohanty is a Postdoctoral Researcher at the Indian Institute of Science (IISc), working in Neuroscience.

LinkedIn: 

Suvendu Kumar https://in.linkedin.com/in/suvendu-kumar-50735717a

Saveena Solanki https://www.linkedin.com/in/saveenasolanki/

Mudit Gupta https://www.linkedin.com/in/mudit-gupta-375a18212/

Sonam Chauhan https://www.linkedin.com/in/sonam-chauhan-a56138170/

Sanjay Kumar Mohanty https://www.linkedin.com/in/sanjay-741mohanty/

Lab PI name: Dr. Gaurav Ahuja

University or institute: Indraprastha Institute of Information Technology, Delhi

Dr. Gaurav Ahuja LinkedIn: https://in.linkedin.com/in/gaurav-ahuja-45a80528

Lab Website  https://www.ahuja-lab.in/

What was the core problem you aimed to solve with this research?

Most molecular artificial intelligence models see a chemical compound through only one lens, such as its structure, physicochemical properties, or biological activity. Each lens captures useful but incomplete information, while combining many data types directly is computationally expensive and often fails when one modality is unavailable. We aimed to create a unified molecular representation that retains complementary information, remains robust to missing source features, and can be generated from a simple SMILES string.

Chemical Dice Integrator Unifies Molecular Data for AI-Driven Drug Discovery
CDI integrates six complementary molecular modalities into CDI-Basic and distills their shared information into CDI-Generalized, which produces unified embeddings directly from canonical SMILES for scalable downstream use.

How did you go about solving this problem?

We developed Chemical Dice Integrator, or CDI, around six complementary molecular modalities: Mordred physicochemical descriptors, GROVER graph features, ImageMol visual features, Signaturizer bioactivity features, MOPAC quantum properties, and ChemBERTa chemical-language features. A hierarchical autoencoder first learns shared information across these views and fuses them into an 8,192-dimensional CDI-Basic representation. We then trained a Mamba state-space model to distill this rich representation directly from canonical SMILES, producing CDI-Generalized. We evaluated the embeddings across 23 classification datasets comprising 171 tasks, 10 regression datasets, scaffold-based out-of-distribution settings, limited-data experiments, modality ablations, retrieval analyses, and external chemical libraries.

“CDI turns complementary molecular knowledge into an accessible representation that can support more reliable and scalable chemical discovery.”  – Saveena Solanki (Corresponding Author)

How would you explain your research outcomes (Key findings) to the non-scientific community?

A molecule can be described in several different ways, much like a person can be described by appearance, behaviour, history, and skills. CDI learns from several such descriptions together and compresses them into one reusable molecular fingerprint. Its distilled version needs only the molecule’s SMILES text, yet retains much of the broader information learned from the six original views. Across diverse prediction tasks, this unified representation was robust and broadly useful, while CDI-Generalized generated valid embeddings for every valid SMILES tested in the evaluated external libraries.

“CDI transforms diverse molecular perspectives into a unified representation, enabling robust predictions and accelerating practical chemical discovery across data-limited settings.” – Prof. Gaurav Ahuja (Corresponding Author)

What are the potential implications of your findings for the field and society?

CDI provides a practical bridge between information-rich molecular modeling and real-world deployment. Researchers can use a single SMILES string to obtain a reusable embedding for toxicity, bioactivity, physicochemical-property prediction, similarity search, and compound prioritization. This could help teams screen large chemical libraries more efficiently and focus laboratory effort on stronger candidates. The method is intended to support, not replace, experimental validation; its societal value will depend on careful prospective testing, transparent reporting, and responsible use in drug discovery.

What was the exciting moment during your research?

The most exciting moment was seeing CDI-Generalized reproduce the information-rich embedding from SMILES alone and maintain complete embedding coverage across chemically diverse external libraries. That result showed that the work had moved beyond a complex multimodal training system into something researchers could realistically deploy, including when upstream quantum or pretrained encoders fail for individual molecules.

Paper reference and project links: Kumar, S., Solanki, S., Gupta, M. et al. Scalable molecular representations enabled by multimodal fusion and sequence distillation. Nat Commun (2026). https://doi.org/10.1038/s41467-026-77700-z .

More news from academia

Related Articles

Subscribe to Newsletter

SciFocus Newsletter by BioPatrika