Highlights
Artificial Intelligence has become remarkably good at understanding human language. But in drug discovery, the language that matters most is not English. It is chemistry.
Every potential medicine begins as a molecular structure. Scientists must accurately interpret atoms, predict properties, assess safety, design experiments, and make critical research decisions. Yet many large AI models, despite their impressive reasoning abilities, can struggle with this seemingly fundamental task. They may generate convincing explanations while silently misinterpreting the molecule itself.
In pharmaceutical R&D, such errors are not academic. A single incorrect molecular interpretation can send researchers down the wrong path, leading to weeks of unnecessary laboratory work and testing.
This is where the next frontier of enterprise AI begins to look very different. In science, the biggest model is not always the best co-scientist. Drug discovery does not need an AI that can merely discuss chemistry — it needs one that can understand it. The future may therefore belong not to ever-larger models, but to models that understand a domain deeply enough to be trusted.
To explore this potential, TCS developed a molecular intelligence model, a domain-tuned chemistry Small Language Model (SLM) built by fine-tuning NVIDIA Nemotron 3.5 to understand molecular structures with higher precision. Instead of relying on scale alone, the model combines the foundation of NVIDIA Nemotron, machine-verified chemistry knowledge and TCS’ deep Life Sciences expertise.
This model demonstrates a stronger understanding of molecular structures than significantly larger general-purpose reasoning models, reinforcing that in drug discovery, domain intelligence can matter more than model size.
Drug discovery depends on making the right scientific decisions early. Before a molecule enters synthesis, screening, or optimisation, researchers need confidence that its structure has been interpreted correctly.
Known as structural fidelity, this ability to accurately interpret atoms, bonds, rings, and functional groups from molecular representations such as SMILES underpins every downstream AI-driven drug discovery task.
On a rigorous RDKit-verified benchmark containing 100 molecular structure questions, TCS’ molecular intelligence model achieved 80% accuracy, outperforming both the Nemotron 3.5 base model at ~76% and a leading general-purpose reasoning model at ~62%. The improvement is not merely statistical; it translates into greater confidence in molecular interpretation, helping researchers make better-informed decisions at the earliest stages of discovery.
The rapid progress of large language models has created a natural assumption that bigger models will continue to deliver better outcomes across every domain. But scientific discovery places different demands on AI.
Compared with a larger frontier reasoning model, TCS’ molecular intelligence model reduced molecular structure-reading errors from approximately 38% to 20%, representing nearly a 47% reduction in error rates. This means fewer molecules enter a discovery workflow carrying silent structural mistakes that can invalidate downstream predictions.
In practical terms, each incorrect molecular interpretation can trigger a wasted 4-8 week synthesis and assay cycle, consuming valuable scientist time and laboratory resources. By preventing these errors earlier in the pipeline, AI can help researchers focus their efforts on the molecules most likely to matter.
More broadly, these findings challenge a common assumption in enterprise AI: larger models are not always the most effective models. In highly specialised domains, performance increasingly depends on how well AI understands the underlying science, context, and workflows it is meant to support.
Accuracy is essential, but it is not enough.
In regulated environments, scientists need AI systems that recognise uncertainty and identify invalid inputs rather than confidently producing incorrect answers.
One of the standout characteristics of TCS’ molecular intelligence model is its ability to detect chemically impossible structures and reject invalid molecular inputs. This calibration is critical for building trust and ensuring reliable decision-making for scientific workflows.
The model also demonstrates good performance in identifying functional groups, understanding molecular scaffolds, stereochemistry, and ring systems. These capabilities form the scientific foundation required for future AI-assisted molecule design and optimisation.
For researchers, this represents a shift from AI as a conversational tool to AI as a scientific collaborator, one that can support discovery with greater precision, reliability, and contextual understanding.
The development of TCS’ molecular intelligence model represents only the first milestone in a larger journey.
The immediate breakthrough lies in helping AI read chemistry more accurately. But the long-term opportunity is much broader: to enable AI systems that can generate new molecules, orchestrate specialised scientific models, and support closed-loop discovery workflows where design, testing, analysis, and learning continuously reinforce each other.
The journey progresses through four stages:
This evolution has the potential to fundamentally reshape how therapies are discovered, developed, and brought to patients.
For pharmaceutical and life sciences organisations, innovation cannot come at the expense of intellectual property protection. Molecular libraries, screening data, and discovery pipelines represent some of the industry’s most valuable assets.
For AI to be trusted in drug discovery, it must not only understand chemistry but also respect the industry's stringent requirements around data privacy, security, compliance, and IP protection.
Built for self-hosted deployment, TCS’ molecular intelligence model allows enterprises to maintain full control of sensitive research data while also benefiting from predictable infrastructure costs. This sovereign-AI approach helps organisations pursue AI-led innovation without compromising data protection, compliance or intellectual property guardrails.
The TCS exploration with NVIDIA Nemotron points to a broader shift in enterprise AI strategy.
For Life Sciences organisations, the opportunity is not simply to adopt the largest available model. It is to build and deploy AI systems that understand the scientific domain, respect enterprise data boundaries, and deliver measurable impact across research workflows.
As AI becomes increasingly embedded across the R&D value chain, competitive advantage will belong to organisations that can combine advanced AI capabilities with deep scientific expertise. Domain-tuned AI can help enterprises:
The broader lesson is clear. In scientific domains, the value of AI will be defined not by how generally intelligent it appears, but by how deeply and reliably it understands the work it is expected to support.
The exploration undertaken by TCS demonstrates that the next frontier in drug discovery AI is not simply larger models, but smarter application of domain knowledge. By combining NVIDIA Nemotron 3.5 with machine-verified chemistry knowledge and deep Life Sciences expertise, TCS has shown how domain intelligence can unlock meaningful gains in scientific accuracy and trust.
The future of drug discovery will not be shaped by AI that knows a little about everything. It will be shaped by AI that understands the science deeply enough to become a trusted research collaborator.
In the next era of scientific AI, competitive advantage will belong not to organisations with access to the biggest models, but to those that can combine AI with deep domain expertise. The ability to teach AI the language of science may ultimately prove more valuable than teaching it everything.
This effort was driven along with Deepesh Aggarwal, Dr. Navneet Bung, Jinal Shah, Aryan Kasat and Vaishnavi M.