On July 7, 2025, appellate lawyer Adam Unikowsky published a provocative claim on Substack: “a robot lawyer would be an above-average Supreme Court advocate.” He had given Claude Opus 4 the briefs and key precedents from Williams v. Reed, a case he had recently argued before the Court, and asked the model to answer the justices’ actual questions. Producing the audio required substantial editing, and Claude sometimes missed what a question was getting at. Still, the result was striking: Many of Claude’s answers sounded plausible, and its emulation (through other software) of Unikowsky’s voice was eerily on point. Justice Elena Kagan later said Claude had done an “exceptional job” of analyzing the case’s difficult Confrontation Clause issue—before adding, “It just seems ridiculous that Claude could do an argument better ... than I could. I kind of think I’m better than Claude.” How should courts respond to models that can produce work of this quality, alongside confident and consequential errors? Last month, we had a chance to explore that question with federal judges, administrators, and attendees at the D.C. Circuit Judicial Conference, where we gave a presentation on the state of artificial intelligence (AI) and the law. The most consistent question we faced from the audience was not about litigant use of AI, but about whether courts themselves should use AI, and if so, how they should use it. At present, judges generally face an unsatisfactory set of choices: (a) abstain from using frontier AI (the most advanced, general-purpose models available at a given time), (b) experiment through personal accounts subject to consumer-facing terms, or (c) rely on the AI products bundled into familiar legal-research platforms.AI has arrived at the courthouse steps—and judges and lawyers are already testing its edges. Three uses now at the frontier of “AI lawyering” and “AI judging” make clear how fast this is moving, and how little institutional infrastructure exists to meet it. What’s missing is a secure, judiciary-controlled environment where judges and court staff can compare leading models, study their failures, and develop informed policy before widespread utilization in live adjudication. The question now is what kind of door to courts’ use of AI gets opened.The New Legal-AI Landscape Two technical developments have helped reshape AI’s role in law in the past year. First, frontier models have become substantially better at sustained legal analysis. For example, on LegalBench, a leading benchmark for legal tasks, the top-performing model now scores nearly 89 percent accuracy across a wide range of legal-reasoning tasks. In a recent, blinded study, contract law professors preferred AI answers to common hypotheticals over answers written by peer professors in roughly 75 percent of comparisons. (They also rated whether answers were harmful and AI models were harmful 3.5 percent of the time; shockingly, one law professor’s answers were deemed harmful nearly 40 percent of the time.) Second, models increasingly operate as agents. They can search the web, navigate databases, write and run code, retrieve documents, remember intermediate results, and revise a research plan as they work. For instance, as we demonstrate in part below, a lawyer could deploy an agent with a single prompt to find a federal rule, analyze thousands of comments in response, and draft an analysis of which agency responses to issues were weakest to prepare a complaint.Concerns, of course, abound. One evaluation, published in the peer-reviewed Journal of Empirical Legal Studies and co-authored by one of us, found that leading AI legal-research products hallucinated between 17 and 33 percent of the time, despite claims that retrieval systems would eliminate fabricated authorities. A growing database now catalogs more than 1,800 legal decisions worldwide discussing alleged or established AI hallucinations. The errors have reached self-represented litigants, lawyers at elite firms, government attorneys, and even judges’ chambers. Two federal judges acknowledged that chambers personnel used AI in drafting orders containing hallucinations, prompting criticism from Senate Judiciary Committee Chair Chuck Grassley. One of the judges had attributed entirely hallucinated quotes to the defendants and indicated that motions to dismiss were denied when they had actually been granted. Careless AI, in the words of Chief Justice John Roberts, risks “dehumanizing the law.” Exhibits From the Frontier of AI Lawyering Conversations about law and AI can drift quickly into the abstract. What cuts through that is watching the technology actually work—where it helps and where it falls short becomes far clearer when tested against real tasks courts face: simulating and preparing for oral argument, researching across sprawling statutory and regulatory corpora, and shaping how judges interpret the law. Oral ArgumentIn his Substack post, Unikowsky showed that a large language model (LLM) could give sophisticated answers to questions actually asked at Supreme Court oral argument that looked and sounded much like his own. We were curious how well this would work with newer models so that judges could better understand AI’s strengths and limits as an advocate and, critically, how well the process could be inverted to help judges themselves prepare for oral argument.To this end, we replicated Unikowsky’s experiment using a D.C. Circuit case—TikTok Inc. v. Garland—and produced similarly impressive answers. For example, on a question asking the government to distinguish Lamont, a central petitioner’s precedent for Americans’ First Amendment right to receive foreign speech, the model’s initial response reached arguments that Justice Department lawyer Daniel Tenny developed only after subsequent pressing and prodding from Judge Sri Srinivasan. The replication, however, also exposed familiar weaknesses: When a question embedded a fabricated holding, the model accepted the premise and reasoned from it, a clear example of model sycophancy. (The question in this instance proposed an incorrect reading of Holder, a key case setting the terms of deference to Congress and the executive when evaluating national security threats.) What about the potential to draft questions to help judges prepare for oral argument? Former Acting Solicitor General Neal Katyal recently reported using AI to prepare for oral argument before the Supreme Court in Learning Resources, Inc. v. Trump. He said that the AI system had accurately predicted several lines of questioning he ultimately received, including a question from Justice Amy Coney Barrett that he described as “almost verbatim” what the system had forecast. To assess this potential for judges, we prompted a model to generate the questions a D.C. Circuit panel might ask after reviewing the briefs and relevant precedents. The questions were plausible enough that, when we showed conference attendees one real question and one generated question, the room was divided over which was human. That said, what is helpful for the solicitor general may be less helpful for judges. Evaluating what amounts to a good question is harder than evaluating a good answer. When generating questions for judges, the model concentrated on issues in proportion to the briefing—in other words, if an issue took up about 50 percent of the briefs, it also appeared in about 50 percent of the questions—while judges often use oral argument to probe what the parties under-briefed or avoided. It also struggled to reproduce the interactive quality of a panel, producing stand-alone questions, when many questions in practice build on or strategically respond to fellow judges. Prompting a model to imitate a particular judge or interpretive methodology like originalism may introduce some variation in proposed questions, but LLM “personas,” such as a synthetic Justice Barrett, currently remain imperfect. These are areas of ongoing research and improvement with much potential, but answering a question is not the same as asking it.Statutory and Regulatory Research Models may be especially useful to judges and clerks when legal research requires searching an enormous body of text. Stanford RegLab developed the Statutory Research Assistant, or STARA, to conduct and document large-scale, systematic surveys. STARA is an automated system that represents code provisions much in the way we teach statutory interpretation, by making the structure, context, definitions, and cross-references clear. This enables an LLM to perform accurate “statutory surveys,” or compilations of all provisions relevant to a particular legal issue, together with detailed annotations and reasoning.We tested this system against Justice Stephen Breyer’s dissent in the 2010 case Free Enterprise Fund v. Public Company Accounting Oversight Board. To illustrate the possible reach of the Court’s holding, which found the board’s dual for-cause removal provision unconstitutional, Justice Breyer produced a 43-page appendix surveying federal offices with similar provisions. In many other cases, the same general question—how broadly a holding sweeps—is important, but difficult to ascertain. The 43-page appendix presumably took his chambers multiple days to search the over 60,000 pages of the U.S. Code for matching provisions and to analyze potential candidates for fit. (When the Department of Justice had attempted a similar task to count the number of federal crimes, the team of government lawyers concluded—after two years of work—that they could only produce a rough guess.) When we asked STARA to complete the same exact task, it—in less than one hour—recovered 93 percent of the heads with for-cause removal provisions and identified at least eight more that plausibly met Breyer’s criteria. One of the provisions STARA had identified (and Breyer had not) was the 2008 statute protecting the director of the new Federal Housing Finance Agency from removal except for cause. That was hardly an obscure edge case: About a decade later, the Supreme Court invalidated that very provision in Collins v. Yellen. STARA wasn’t perfect, but neither were the clerks. It produced in under an hour what likely took the justice’s chambers days. Such a tool could enable judges to understand the implications of a ruling when such resources would not be available. Recent frontier models can take surprisingly good first stabs at otherwise daunting questions without custom infrastructure. We demonstrated that possibility by asking Claude Fable 5 to analyze the Council on Environmental Quality’s (CEQ’s) rescission of its National Environmental Policy Act regulations. To complete this task, the agent had to determine whether commenters raised objections based on statutory authority or reliance interests and whether CEQ answered them adequately. You can watch the somewhat dizzying screen recording below.In 18 minutes, the agent located the docket (which had a staggering 88,800 comments), devised a search strategy, downloaded and parsed more than 500,000 words (or approximately 1,200 pages) of comments, extracted supporting quotations, and produced a structured memorandum with citations and verifiable pin-citations. The resulting memo was a strikingly useful starting point—although of course incomplete. On careful review of the agent’s approach and tool calls, we saw that it had retrieved only the top 10 results for each search, apparently conserving cost, but potentially missing many relevant comments. Used in practice for similar tasks, a judge evaluating only a polished memorandum might see extraordinary speed and apparent comprehensiveness, but inspection of the agent’s process also reveals the retrieval choices that bounded its answer. For example, the LLM might have arbitrarily decided on a search strategy that queried for “law professors” but not “Sierra Club,” and missed a key, distinct issue raised only by the Sierra Club that for some reason failed other search terms. The judge then might have an incomplete—without knowing necessarily whether or how it is incomplete—sense of the distinct issues raised that may or may not necessitate agency response under the Administrative Procedure Act.Interpretation The most controversial use involves interpretation of the law itself. Judge Kevin Newsom of the U.S. Court of Appeals for the Eleventh Circuit has argued that LLMs might help judges triangulate the ordinary meaning of legal text. The extent to which LLMs are well-suited to this task is hotly debated. On the one hand, a recent experiment—in which a researcher asked lawyers, non-lawyers, and LLMs to interpret the plain meaning of words at issue in real cases—found that LLMs performed comparably to judges in predicting what the “consensus” ordinary meaning of a phrase was. On the other hand, new research shows that LLM judgments about ordinary meaning are extremely sensitive to exact prompt phrasing, suggesting that this process is easily manipulable. Other scholars identify a deeper institutional concern. Model outputs reflect technology firms’ undisclosed choices about training data, fine-tuning, and system design. They may overrepresent elite or foreign usage of terms, as was identified in the case of ChatGPT’s overuse of the word “delve” because much of the firm’s posttraining data was collected from Nigerian and Kenyan data workers who use the word much more frequently.If judges treat an LLM as a neutral measure of ordinary meaning, it could quietly transfer interpretive influence to model developers. If U.S. leaders would not look favorably on a judge explicitly using Bismarck’s constitution to interpret the U.S. Constitution, they should similarly not be so kind to the implicit (and unknown) use through an LLM. Courts need to actively test models across prompts, model families, versions, and human baselines to discover these quirks, adjust their usage accordingly, and know how and where to scrutinize litigant arguments that may be leveraging LLMs as interpretive tools. An AI Experimentation “Playground” for the Courts Courts are confronting generative AI without a settled institutional strategy. A 2026 survey of 112 federal judges found more than 60 percent of them had used at least one AI tool in their work, although only 22 percent used AI weekly or daily. About 20 percent formally prohibited AI in chambers, 18 percent discouraged it, and nearly one-quarter of judge participants had no official policy. Critically, judges were more likely to use legal-specific AI products—think AI as integrated into existing legal providers, such as Lexis and Westlaw—than general-purpose systems like Claude or ChatGPT. The federal judiciary has made some initial moves toward addressing AI use in chambers. An internal task force within the Administrative Office of the U.S. Courts issued nonpublic interim guidance in July 2025 addressing the use, procurement, and security of AI tools, while cautioning against delegating core judicial functions. But restrictions alone cannot supply the understanding of AI that courts need. Unless barred, litigants will almost certainly use frontier systems to conduct research, review records, develop arguments, and generate filings. As a result, judges will increasingly have to evaluate claims about what those systems can do, whether their use has tainted evidence or legal work, and what discovery concerning them should entail. Courts will struggle to perform that function without concretely understanding the promise and failure points of systems. We consider several possible futures for how U.S. courts will approach AI. Banning AIA ban on judicial AI, adopted by some chambers, can reduce immediate risks, particularly when courts lack secure systems or adequate training. California’s statewide rule, for example, expressly allows an individual court to prohibit generative AI from being used in chambers. But a prohibition on AI in chambers would just allow the information asymmetry between courts and litigants to grow, since litigants will certainly continue to use AI, and courts that have not been experimenting with the tool will be ill-equipped to detect and scrutinize the types of decisions that AI tools increasingly make (for example, what search terms it used in querying large document sets, or what source it weighed the most heavily when determining ordinary meaning). A ban on AI use in court systems may also not prevent a curious judge or clerk from accessing a model on a personal device, nor is it clear how such bans would affect chambers’ use of Google search, which increasingly relies on AI overviews in providing answers.Laissez FaireAt the opposite extreme, some may see little reason for institutional involvement regulating who can and can’t use AI in the judiciary. The belief here is that judges should just decide by themselves to use whatever AI product they deem suitable. To avoid confidentiality risks, a judge working exclusively with public materials can use a personal account and disable model training so that future models are not trained on the judge’s unfinished, deliberative work product. But merely disabling training does not eliminate retention of data for human abuse review and account-level personalization, or later changes to product terms. Judicial deliberations deserve protections defined by enforceable contracts and court policy rather than a toggle whose scope an individual user may misunderstand. Safeguards, such as contractual agreements to ensure zero data retention, should ensure that technology firms do not inadvertently develop databases filled with potentially nonpublic information from litigants as well as judges’ deliberative processes. An AI firm before a court should not be able to revise its litigation strategy based on that judge’s prompt history, available only to that firm.Judges also shouldn’t have to hope that a vendor-controlled interface will offer judicial users the same product offered to other users. Uber offers an instructive analogy: The company identified police regulators and served them a simulated version of its platform populated by “phantom cars” that never arrived.Given the growing range of cases involving AI companies, the judiciary cannot afford to be shepherded off to a sanitized version of an LLM lacking features or risks that might be at issue in a case. Officials should be able to minimize the system’s ability to identify and selectively respond to them by negotiating contracts that limit vendors’ abilities to differentiate their product based on their judicial status. Existing ProvidersThe more likely status quo is AI use through familiar legal providers such as Westlaw and Lexis, which judges are already most comfortable with, according to the recent survey. This approach is advantageous because these platforms already have access to extensive legal corpora, and because so many lawyers and judges are already familiar with using these systems.Despite these advantages, exclusive reliance on status quo legal research tools would be a fatal mistake. The capabilities demonstrated above arise from agentic models that can plan, browse, download, write code, revise search strategies, and work across heterogeneous sources. These capabilities have historically reached general-purpose frontier platforms before being incorporated into specialized legal products. In our research, Westlaw and Lexis AI tools have typically significantly underperformed on legal tasks. For example, when tasked to identify all federal crimes in the U.S. Code, Westlaw’s AI Jurisdictional Survey found less than half as many provisions as Gemini Deep Research, and hallucinated 121 statutory provisions. In more recent work classifying state unemployment insurance provisions, Westlaw and Lexis fared even worse than standard retrieval augmented generation models, by 4 to 39 percent on a standard measure (F1) that accounts for misses and false hits. And even in a straightforward query comparing several states’ open meeting laws conducted this month using Westlaw’s Jurisdictional Survey, the tool analyzed Pennsylvania’s municipal-pension funding law rather than its Sunshine Act (which mandates agencies to deliberate in open and public meetings), did not even provide an answer for Tennessee (which does have an Open Meetings Act), and substituted features of Texas’s Election Code for its open meeting laws. Given the same task, Claude Fable 5 certainly made mistakes, but its answer was substantially more complete and accurate on at least seven of 10 states compared to Westlaw’s. (Claude did analyze Tennessee’s Open Meetings Act, Pennsylvania’s Sunshine Act, and Texas’s open meeting laws, and quite accurately for all three, with the exception of missing Tennessee’s penalties for infractions. Unfortunately, LexisNexis and Thomson Reuters opted out of participating in other efforts to benchmark performance across providers.) Products vary, especially at the fierce pace of AI innovation. And this is exactly the kind of comparison that courts need to see and conduct firsthand. On-Premise AdoptionThe most ambitious path toward integrating and regulating AI use in chambers could be an on-premise system that would provide strong control over data and customization. The South Korean Supreme Court is developing its own large language model specialized for sentencing trained on internal court data. Such development, which cost South Korea more than $7 million, is likely beyond the IT capacity of U.S. courts, which continue to struggle with staffing shortages and with more basic tasks like document management. Even if the U.S. did have these capabilities, developing on-premise models means relying on smaller, open models rather than frontier models like those from OpenAI and Anthropic. While the performance gap between open models and frontier models appears to be closing, such as with the launch of Kimi K3, the gap’s future is uncertain. An AI PlaygroundThe near-term solution that courts should take seriously is a secure, vendor-neutral AI sandbox, or a controlled environment where users can experiment with various models under clear data privacy guarantees. The judiciary could operate the interface, authentication, policy controls, and evaluation records while accessing cloud-hosted models through centrally negotiated enterprise or API agreements. Stanford, among other institutions within and outside government, has created an AI “playground,” or a centralized environment where users can compare models from OpenAI, Google, Anthropic, and other providers. This allows faculty, students, and staff to experiment with prompting various models at no cost to themselves under an institution-wide agreement that ensures that uploaded files are not shared externally or used to train models—exactly the kinds of guarantees the judiciary deserves. According to our discussion with the university’s senior director for enterprise AI, the platform was built by lightly modifying an open-source interface (LibreChat) and initially deployed without a dedicated AI team. Current operation requires one to two engineering full-time employees, one part-time employee, and some contractor support, with model costs on the order of tens of thousands of dollars per month across the university’s nearly 40,000 students, faculty, and staff. A judicial version of this playground could begin with permission only to consult public records, and should offer access to several frontier models. Judges could submit the same research question to different models, inspect their sources and intermediate steps, and compare performance against conventional research and known outcomes. The system should ideally record the model and version used, since results produced today may disappear after an update next month. The governing agreements—like Stanford’s—should prohibit training, require zero data retention, restrict provider personnel’s access to prompts, prohibit account-level advertising or behavioral profiling, and require notice of material model changes and security incidents. The courts would also benefit from a collective acquisition approach—that is, working as a judiciary-wide system rather than court by court or circuit by circuit—to negotiate with AI vendors, who are accustomed to working with, and offering discounts to, large-scale purchasers. (Note that even a judiciary-wide contract’s purchase commitments may pale in size compared to those of a mid-sized company with engineers.) Equally important to procurement arrangements is the collective learning environment that the courts should foster. While general AI training can have some value, engagement with AI should not be left to hypotheticals—there is no substitute for seeing an agent in action to meaningfully understand the technology’s capabilities and limitations. To facilitate shared learning, each court could hold periodic sessions in which judges and clerks share useful workflows, unexpected failures, prompt sensitivities, and verification costs. Failures and successes discovered in one’s chambers should become institutional knowledge. The judiciary does not need to choose today which legal tasks AI should ultimately perform. It does need the capacity to investigate that question independently, securely, and with direct access to the most capable models. A judicial AI playground would preserve caution while preventing that caution from hardening into technological inertia or institutional ignorance.