What is the purpose of the CiteClear benchmark?
The CiteClear benchmark is designed to measure and publicly report on the reliability of AI-generated legal citations. Unlike proprietary benchmarks run by AI vendors themselves, our methodology is transparent, our prompts are published, and our results are independently verifiable. We aim to provide objective data that helps legal professionals understand both the capabilities and limitations of AI tools for legal research.
How does CiteClear's methodology differ from other AI benchmarking efforts?
CiteClear focuses specifically on legal citations rather than general factual accuracy. We use a deterministic-first approach that doesn't rely on the AI models we're evaluating. Our prompts are designed to test real-world legal research scenarios, not synthetic benchmarks. Most importantly, we publish our entire methodology, allowing anyone to reproduce or audit our work.
Can I contribute to or reproduce the CiteClear benchmark?
Yes. Our methodology and sample prompts are fully published. Researchers, journalists, and developers are encouraged to reproduce our tests, extend our methodology, or contribute new prompt sets. All our benchmark materials are available for independent use. Contact us at the email in our footer for collaboration opportunities or access to our full prompt library.
Why focus only on citations rather than full legal analysis?
Citations are the most verifiable component of legal writing. Unlike legal analysis, which requires subjective interpretation of complex legal principles, a citation either exists or it doesn't. This makes citations an ideal objective measure for AI reliability. Additionally, focusing on citations allows for privacy-preserving verification without sharing confidential client documents, which is essential for legal work.
Does CiteClear sell benchmark data or reports?
No. Our benchmark results are published freely on this page. We do not sell data, reports, or access to proprietary benchmark information. Our goal is transparency in AI legal citation reliability, not commercialization of benchmark data. All results are available to everyone at no cost.
How often will CiteClear update the benchmark results?
We plan to run benchmark cycles on a regular basis, particularly when new AI models are released or when existing models are significantly updated. We'll publish new results as they become available. The methodology will remain consistent across runs to ensure comparability, though we may add new prompt categories over time.
What AI models does CiteClear test?
We test widely available commercial AI models, including but not limited to: ChatGPT (various versions), Claude (various versions), Gemini, Perplexity, and other models that are accessible through standard APIs or consumer interfaces. We document which specific model versions we test in each benchmark cycle.
How does CiteClear handle model updates and version changes?
AI models are frequently updated, which can change their performance. For each benchmark run, we clearly document: 1) The specific model name and version, 2) The date of testing, 3) Any known model updates between runs. This allows for fair comparison across different benchmark cycles and helps track performance over time.
What about proprietary or specialized legal AI models?
Currently, our benchmark focuses on general-purpose AI models that are widely accessible. We may expand to include specialized legal AI tools in the future. If you're a legal AI vendor and would like your product included in our benchmark, please contact us. We maintain the same transparency requirements for all tested systems.
Can I use CiteClear's benchmark for my own research?
Yes. Our methodology and sample prompts are published for anyone to use. We only ask that you: 1) Cite CiteClear as the source of the methodology when publishing results, 2) Make your own results and methodology transparent, 3) Don't use our work for commercial benchmarking products without permission. Academic and journalistic use is encouraged.