Data Fusion & Analytics
Scale Evaluation (SEAL)
Independent test, evaluation, and red-teaming for frontier AI models.
Scale Evaluation, run by Scale's Safety, Evaluations and Alignment Lab (SEAL), provides model test-and-evaluation, benchmarking (the SEAL Leaderboards), and adversarial red-teaming using curated private datasets and verified domain experts. Its red-teaming product was selected by the White House for public assessments of leading AI models, and Scale was tapped to set the Pentagon/CDAO's path for testing and evaluating LLMs.
What it does
- ▸Frontier-model benchmarking (SEAL Leaderboards)
- ▸Adversarial robustness and red-team testing
- ▸Tamper-proof private evaluation datasets
- ▸Government T&E of LLMs for the DoD
- ▸Safety and alignment research
Where it's deployed
- ◍The White House
- ◍Pentagon / CDAO LLM test & evaluation
Agencies using it
DoD