Put competing AI systems through the same real tasks and discover why “best” depends on what you measure.
Design a fair comparative evaluation of several AI models or configurations for a narrow, benign use case using transparent criteria and repeatable tests.
Original Compass project by Davisville LabsReviewed August 2026Goal real work, evidence, and reflectionEditorial standards →
Why this project matters
Learn by doing something real.
Design a fair comparative evaluation of several AI models or configurations for a narrow, benign use case using transparent criteria and repeatable tests. This project feels different from a school assignment because the student serves a team choosing an AI tool for a defined low-risk workflow, creates something visible, tests it honestly, responds to feedback, and explains the decisions behind the final result.
What you will create
Your project plan
Use these deliverables as milestones for Run an AI Model Bake-Off. Adapt the details to your interests, available time, and real-world opportunities.
01
User, data, and risk requirements
Research use case, test set, scoring rubric, blind review, latency, cost, consistency, hallucination, privacy, accessibility, and model updates. Document the needs, constraints, credible sources, and perspectives of a team choosing an AI tool for a defined low-risk workflow.
02
Technical architecture and test plan
Turn the evidence into a focused plan for the comparative AI benchmark and recommendation, including success criteria, ethical boundaries, scope, and a realistic path to completion.
03
Working product or prototype
Create the first complete version of the comparative AI benchmark and recommendation. Include a use-case specification, representative test set, scoring guide, model run log, blinded outputs, quantitative results, qualitative error analysis, and recommendation.
04
Reliability and usability test log
repeat a sample across days or settings, use multiple reviewers for subjective scores, and report whether small changes reverse the ranking Record what happened, what failed, what users or reviewers said, and which changes you made.
05
Final technology case study
Publish the final comparative AI benchmark and recommendation with a portfolio-ready case study showing the challenge, evidence, process, revisions, results, limits, and next version.
Make the work stronger
What separates a finished project from a meaningful one?
For Run an AI Model Bake-Off, use these checkpoints to protect the quality of the work without turning the project into a performance for admissions.
Avoid this
Starting the comparative ai benchmark and recommendation before understanding a team choosing an AI tool for a defined low-risk workflow
Treating assumptions or internet opinions as real evidence
Choosing a scope too large to test and finish with care
Evidence that it is working
The audience and real problem are clearly defined
Research and evidence are documented
A working comparative AI benchmark and recommendation was created
If you want to go further
Test the work with at least five additional people from a team choosing an AI tool for a defined low-risk workflow
Interview a professional connected to AI Product Management and compare their advice with your approach
Skills you can build
Capabilities that travel beyond this project.
Benchmark DesignQuantitative EvaluationAI Product AnalysisTechnical Problem SolvingResponsible Technology
College application value
Use the project as evidence, not decoration.
Can demonstrate initiative, curiosity, reflection, and growth through the student’s choices, response to setbacks, work with a team choosing an AI tool for a defined low-risk workflow, and development of Benchmark Design, Quantitative Evaluation, and AI Product Analysis.
The goal of Run an AI Model Bake-Off is not to manufacture an impressive activity. Do real work, keep evidence of the process, and reflect honestly on what changed.
Related college majors
Which fields connect to this work?
Use Run an AI Model Bake-Off as a clue, then open a related major guide to compare coursework, career directions, reality checks, and other ways to test the field.
Run an AI Model Bake-Off is an original Compass educational starting point designed around real work, visible evidence, feedback, and reflection. It is not an admissions guarantee.