The Supreme Court Just Caught AI Faking Precedents. The Harder Problem Is What It Cannot Catch.

Mariya Noor

On 2 July 2026, the Supreme Court of India delivered a judgment unlike any before it. In Pooja Ramesh Singh v. Jammu and Kashmir Bank, it set aside orders of the NCLT and the appellate tribunal after discovering that the tribunal had relied on judgments that did not exist, fabricated by an artificial intelligence tool and cited as if they were law. The Court declared zero tolerance, held that a decision resting on hallucinated material is no decision in the eyes of the law, and directed the Bar Council of India to frame guidelines and disciplinary consequences. Justice Narasimha’s bench insisted that a human must remain in the loop at every stage of adjudication. 

It was necessary, and long overdue. The warning signs had been building for three years. In March 2023, a judge of the Punjab and Haryana High Court consulted ChatGPT during a bail hearing, noting carefully that it was only for a wider picture of the jurisprudence. In August 2023, the Delhi High Court rejected lawyers’ attempts to rely on ChatGPT outputs in the Christian Louboutin trademark case, warning of fictional case laws.

In May 2024, the Manipur High Court itself turned to ChatGPT for research when a police affidavit proved inadequate. A Bengaluru tax tribunal recalled an order involving roughly six hundred and sixty nine crore rupees after four non-existent citations surfaced in it, and the Bombay High Court quashed an assessment built on three invented precedents. AI did not knock on the judiciary’s front door. It came in through the pleadings, the research, and eventually the orders themselves. 

So is the problem solved? Not quite. Fake citations are the failure we can see. A second failure mode remains untouched, and it is harder to catch because nothing about it looks wrong. Both failures grow from the same root, models trained on vast amounts of text that absorb whatever that text contains, whether invented cases or inherited prejudice. 

Consider an experiment by researchers at Stanford. They gave large language models, the technology behind ChatGPT and its rivals, identical prompts and changed only one detail, the person’s name. Names associated with different racial groups received measurably different treatment, even from the most advanced models. The facts never changed. Only the name did, and the answer moved. 

Now bring that home. India is a country where a name carries enormous social information. A first name can signal religion. A surname can signal caste. An address can signal class. If a model shifts its answers over names in American experiments, we should ask what it does when the name is Mohammad Irfan instead of Vikram Rathore, and the question is bail. 

For religion, an early Indian answer already exists. In 2023, researchers from IIT Madras and IIIT Hyderabad tested a machine learning model trained on Hindi bail records from Uttar Pradesh courts. When they swapped Hindu and Muslim names in otherwise similar cases, the model’s predictions changed.

The bias did not run in one simple direction, it varied with the type of case, but it was there, absorbed silently from the data. Caste and class have not yet been tested this way in an Indian legal setting. That is not a reason for comfort. It means we know these tools can pick up one form of social prejudice from Indian data, and we have simply not looked for the others. The researchers ended their paper asking for far more work on fairness in Indian legal AI. That request has gone largely unanswered. 

Here is why bias is the harder problem. A fabricated citation carries the seeds of its own exposure. Anyone can look up a case and find it does not exist, which is exactly how the NCLT order unravelled. A biased output has no such flaw. It cites real cases. It reads fluently.

It does not reveal itself in the ordinary course of litigation, because no litigant ever sees what the tool would have produced for a different name. It shows up only when someone deliberately tests for it, the way the Stanford and IIT Madras researchers did. Until someone does, unequal treatment sits wrapped in the language of neutral technology, and machines are trusted precisely because we assume they have no prejudices to hide. 

Some will say human judges carry biases too, and they do. But a judge’s reasoning is spoken in open court, recorded, and open to appeal. A model’s leanings are buried in billions of parameters that even its makers cannot fully explain. Centuries of legal machinery exist to challenge human discrimination. Nothing comparable yet exists for the machines. 

The pressure to adopt these tools will only grow, and for understandable reasons. India carries more than five crore pending cases, and efficiency is a genuine need, not a villain. But efficiency arguments tend to crowd out fairness questions, because delay is visible and bias is not. A case stuck for ten years makes headlines. A skewed recommendation never does. 

The Supreme Court’s judgment and the Bar Council committee it ordered are aimed at fabrication, and rightly so. But the response should not stop where visibility ends. Three additions would cost little and protect much. First, bias testing before any AI tool is marketed for legal work, using the method researchers have already proven, identical cases, varied identities, and measured outcomes. Second, a disclosure duty, so courts know when AI helped produce what is before them. Third, a designated body with the power to audit these tools as they spread through India’s law offices. 

Pooja Ramesh Singh will be remembered as the moment India’s highest court drew a line against AI’s fabrications. The next line must be drawn against something quieter, the possibility that these tools do not treat everyone equally. The first failure was caught because citations can be checked. The second will be caught only if someone goes looking. That search should begin now, while the rules for AI in Indian justice are still being written. 

References

Girhepuje, S., Goel, A., Krishnan, G. S., Goyal, S., Pandey, S., Kumaraguru, P., &
Ravindran, B. (2023). Are models trained on Indian legal data fair? arXiv.
https://arxiv.org/abs/2303.07247

Haim, A., Salinas, A., & Nyarko, J. (2024). What’s in a name? Auditing large language
models for race and gender bias. arXiv. https://arxiv.org/abs/2402.14875

iPleaders. (2026). AI-hallucinated case law: How fake citations are getting lawyers
sanctioned in India. https://blog.ipleaders.in/ai-hallucinated-case-law-fake-citations-india/

Jaswinder Singh v. State of Punjab, CRM-M-22496-2022 (Punjab and Haryana High Court,
27 March 2023).

Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023)

Md Zakir Hussain v. State of Manipur (Manipur High Court, 23 May 2024).

Pooja Ramesh Singh v. Jammu and Kashmir Bank Ltd., 2026 INSC 668 (Supreme Court of
India, 2 July 2026).

About the Contributor: Mariya Noor is a Legal Tech Associate with three years of experience in the legal technology industry and a fellow of the Public Policy Youth Fellowship 2.0 at IMPRI Impact and Policy Research Institute. Her research examines bias and fairness in artificial intelligence tools entering the Indian legal system. She holds a B.A. LL.B. (Hons.) from Aligarh Muslim University, Malappuram Centre, Kerala.

Disclaimer: All views expressed in the article belong solely to the author and not necessarily to the organisation.

Read more at IMPRI:

Rural Women’s Property Revolution : NHFS-6 on Women Property Ownership

Thirteen years after Delhi gangrape, women’s safety is still a huge issue

Acknowledgement: This article was posted by Vishal Kumar, a Research and Editorial Intern at IMPRI.

Author

Talk to Us