Advancing Responsible, Scalable, and Interpretable AI across Core ML, Law, and Healthcare
Research Fellow at University of Birmingham Dubai | Core AI/ML | LLMs | Legal AI | Healthcare AI | Explainable AI | Multilingual AI
Dr. Shubham Kumar Nigam is a researcher advancing AI memory management, distributed training, test-time training, efficient LLM systems, reasoning, agents, and evaluation—alongside specialized work in Legal AI and Healthcare AI. His work focuses on building transparent, factual, and responsible AI systems.

Recent honours
- Best Paper Award – Bridge between AI and Law @ AAAI 2026
- DAAD Postdoc-NeT-AI Fellow
Research Areas
Exploring the frontiers of AI for Law, Healthcare, and Society
Legal AI
Building AI systems for legal judgment prediction, explanation, retrieval, rhetorical role segmentation, and document generation in the Indian and global legal contexts.
Core AI/ML
Advancing foundational AI and machine learning techniques including model architectures, optimization, and learning paradigms for next-generation intelligent systems.
Evaluation & Benchmarking
Designing rigorous evaluation frameworks, benchmarks, and metrics to assess AI system performance, fairness, and reliability across diverse tasks.
Explainable AI
Developing transparent and interpretable AI systems that provide human-understandable reasoning, mechanistic interpretability, and justification for predictions.
Multilingual AI
Developing AI technologies that work across Arabic, English, Indic languages, and other multilingual settings with cross-lingual transfer capabilities.
LLM Reasoning & Agents
Exploring reasoning capabilities, agentic workflows, tool use, and multi-step problem solving in large language models for complex domain tasks.
Healthcare AI
Creating AI systems for medical dialogue, clinical decision support, predictive modeling, and patient-centric care across multilingual settings.
AI Memory Management
Developing token-efficient memory systems for large language models that compress conversation history without losing critical context and enable extended reasoning.
Distributed Training
Researching scalable distributed training paradigms and efficient model parallelism strategies to train large AI systems across heterogeneous compute clusters.
Test-Time Training
Investigating test-time training and adaptation techniques that enable language models to specialize to unseen domains at inference without expensive fine-tuning.
Featured Publications
IndicMedDialog: A Parallel Multi-Turn Medical Dialogue Dataset for Accessible Healthcare in Indic Languages
Shubham Kumar Nigam, Suparnojit Sarkar, and Piyush Patel
Introduces a parallel multi-turn medical dialogue dataset for accessible healthcare in Indic languages, enabling multilingual medical dialogue systems.
MedAidDialog: A Multilingual Multi-Turn Medical Dialogue Dataset for Accessible Healthcare
Shubham Kumar Nigam, Suparnojit Sarkar, and Piyush Patel
Presents a multilingual medical dialogue dataset designed to support accessible healthcare through AI-powered dialogue systems.
NyayaMind: A Framework for Transparent Legal Reasoning and Judgment Prediction in the Indian Legal System
Parjanya Aditya Shukla, Shubham Kumar Nigam, Debtanu Datta, Balaramamahanthi Deepak Patnaik, Noel Shallum, Pradeep Reddy Vanga, Saptarshi Ghosh, and Arnab Bhattacharya
Proposes NyayaMind, a framework for transparent legal reasoning and judgment prediction tailored to the Indian legal system.
Structure-Aware Agentic and Reinforcement Learning for Legal Judgment Prediction and Explanation
Shubham Kumar Nigam
Explores structure-aware agentic and reinforcement learning approaches for legal judgment prediction and explanation generation.
Featured Projects
KAMAL Health
2 papersProblem: Healthcare systems in multilingual regions lack interpretable AI tools for clinical decision support and patient-centric care.
Method: Developing responsible AI systems that support clinical decision-making across Arabic and English languages with explainable outputs.
Impact: Improves patient care accessibility and clinical decision quality in multilingual healthcare environments.
NyayaRAG
1 paperProblem: Existing LJP systems ignore statutory provisions and judicial precedents, core elements of common law reasoning.
Method: RAG framework integrating case facts, statutes, and semantically retrieved precedents for realistic legal judgment prediction.
Impact: Significantly improves predictive accuracy and explanation quality by grounding predictions in external legal knowledge.
TathyaNyaya / FactLegalLlama
1 paperProblem: Prior LJP datasets use complete judgments including reasoning, unlike real-world early-stage decision-making based only on facts.
Method: Created TathyaNyaya dataset focusing on factual statements, and FactLegalLlama, an instruction-tuned LLaMa-3-8B for fact-based prediction and explanation.
Impact: Enables more realistic legal prediction scenarios and improves transparency in AI-assisted legal analysis.
NyayaAnumana / InLegalLLaMA
1 paperProblem: Existing Indian legal datasets lack scale, diversity across court levels, and comprehensive coverage.
Method: Compiled NyayaAnumana (702,945 cases) and developed INLegalLlama through continual pretraining and supervised fine-tuning on Indian legal documents.
Impact: Achieves ~90% F1-score, setting a new benchmark for Indian legal judgment prediction.
Recent News
Highlights from research, awards, and academic milestones
Best Paper Award at AAAI 2026 Bridge between AI and Law
Received Best Paper Award for work on legal AI.
Invited Talk at COREQ Research Seminar, University of Birmingham Dubai
Presented on building responsible and explainable AI systems for society.
Started as Research Fellow at University of Birmingham Dubai
Leading the KAMAL Health Project on interpretable AI for healthcare.
Four papers accepted at AACL-IJCNLP, NAACL, and COLING 2025
NyayaRAG, TathyaNyaya, LegalSeg, and NyayaAnumana accepted at top venues.
PredEx published at ACL 2024 Findings
Legal Judgment Reimagined: PredEx and the Rise of Intelligent AI Interpretation in Indian Courts.
Rethinking LJP in Realistic Scenarios published at NLLP 2024
Investigated LLM performance in realistic legal judgment prediction settings.
Nonet achieves 1st place in SemEval-2023 Task 6 (Task-C2)
Court Judgment Prediction with Explanation.
ILDC for CJPE published at ACL-IJCNLP 2021
Pioneered the Court Judgment Prediction and Explanation task.
Interested in Collaboration?
I am always open to research collaborations, PhD supervision discussions, invited talks, and industry partnerships in Core AI/ML, Legal AI, and Healthcare AI.