Expert evaluators for AI research. Priced for research.
Preference data (RLHF), model evaluation, benchmarking, expert demonstrations, and red-teaming for university studies, staffed by graduate-level domain specialists from our European expert network. The quality of a lab vendor, at rates built for a research budget.
Your study design and data stay yours. We never train on or resell what the cohort produces. EU-resident evaluators, GDPR-compliant, paid the full rate.
Some studies need scale. Yours needs judgement.
Preference rankings from people who know the field. Rubric grades from professionals. Demonstrations written by someone who has done the job. The program is built for the expert human data behind modern AI, the judgement that trains, aligns, evaluates, and benchmarks models, and when that judgement has to come from an expert, it gives you a staffed, managed cohort matched to your study.
University research groups
Faculty, postdocs, and PhD students whose AI studies need expert human data: preference and reward-model data (RLHF), model evaluation and benchmarking, demonstrations, red-teaming. Open to university-affiliated researchers worldwide.
A staffed expert cohort
Graduate-level specialists matched to your domain and languages, trained on your protocol, managed end to end, with quality review built in. You keep full design control of the study.
Research rates
An hourly expert rate plus a flat service fee, the same two-part shape as the participant platforms you already budget for. One fixed quote before anything starts. Full breakdown below.
The same network that grades EuroExec.
Sovrano Scholars are graduate-level specialists in finance, law, medicine, engineering, and the sciences, recruited from universities across Europe. Every Scholar completes structured training before touching production work, and every project runs with rubric-graded quality review, the same authoring and dual-review method behind the EuroExec benchmark.
The network is EU-resident and engagements run under European data-protection frameworks: a DPA is available on request, and the setup slots directly into a Horizon Europe or ERC data-management plan. Native-language coverage spans the major European languages, with jurisdiction-specific expertise to match.
Specialized in expert human data for AI.
The cohort's specialty is expert human data for AI: if it trains, aligns, evaluates, or benchmarks a model, it is in scope. The most common study shapes:
Pairwise & ranking judgements
RLHF-style preference comparisons over model outputs, ranked by experts who know the domain the outputs are about.
Rubric grading & benchmarks
Model answers scored against your rubric or benchmark, with per-item grades, rater IDs, and agreement statistics you can report.
Gold answers & traces
Expert-written reference answers, worked examples, and reasoning traces for supervised fine-tuning or gold sets.
Adversarial & error analysis
Structured probing by domain experts, with failures labeled against your taxonomy.
Cross-market judgement
Native-speaker evaluation across European languages and jurisdictions, from one coordinated cohort.
Your own protocol
Bring the design your study needs. Your evaluation interface or a shared one we set up, either works.
You send the protocol. We staff it. Data comes back audit-ready.
The study stays yours. We run the part between your task spec and your dataset.
Your tasks, your format
Send tasks in the exact format you would hand an in-house evaluator: instructions, examples, and edge cases. We train the cohort on your protocol as written; the design stays under your control.
We match the cohort
We staff from the Scholar network by domain, language, and availability, and you approve the cohort size, timeline, and fixed quote before work begins.
Data with the paper trail
Completed data with pseudonymous rater IDs, inter-rater agreement statistics, versioned guidelines, and a methods paragraph you can adapt for the paper.
Data you can put in a paper, not just a spreadsheet.
Every engagement returns the dataset plus the paper trail reviewers ask for: pseudonymous rater IDs, inter-rater agreement, versioned guidelines, and a methods paragraph you can adapt. A sample of each:
Preference judgements were collected from N graduate-level evaluators recruited and managed through the Sovrano Research Program. Each pair was independently rated by two evaluators against a shared rubric (Krippendorff's alpha = 0.78). Evaluators were domain specialists in [field], EU-resident, and compensated at their full hourly rate. Guidelines (v3) and per-item ratings are released with this paper.
{
"item_id": "pref_0421",
"task": "pairwise_preference",
"domain": "clinical_medicine",
"rater_ids": ["ev_17", "ev_42"],
"choice": "A",
"agreement": true,
"confidence": [4, 5],
"rationale": "Option A gives the correct first-line therapy; B omits the contraindication.",
"guidelines_version": "v3"
} Illustrative only. Your schema, fields, and rubric are whatever the study needs; this is the shape most groups export.
One band. One fixed quote. The same math we see.
Priced the way participant platforms price, in two parts you can see: the expert rate, €13 to €17 per hour set by the domain, the review depth, the timeline, and the volume, plus a flat 33% service fee that covers training on your protocol, quality review, and project management. The experts receive the full hourly rate. Billed on delivered, QA-passed hours, invoiced in euros as a single purchase-order-friendly invoice.
A full 50-expert study, one 4-hour task each, starts at €3,458 all-in at the base rate: a line item that fits inside a grant's data budget. The €2,600 goes to the experts.
| What sets your rate | Base of the band | Top of the band |
|---|---|---|
| Domain | General graduate-level domains | Regulated specialists: clinical, legal, finance |
| Review depth | Single expert per item, rubric-graded QA | Dual independent review with agreement statistics |
| Timeline | Standard windows, two to six weeks | Compressed timelines and priority staffing |
| Volume | Larger commitments pull toward the base | Small one-off runs sit higher in the band |
The four drivers set the expert rate only. The service fee is flat: 33% on every project, at every size.
| Engagement | Scope | Price |
|---|---|---|
| Pilot | Send 5 tasks our way, the exact format you'd give an in-house evaluator. Completed data back within two weeks. No commitment. | Free |
| Research Pack | 200 expert hours, the standard unit for a study-sized annotation run. Fixed price, single invoice, straightforward to procure. | from €3,458 all-in |
| Full Study | 500+ expert hours: larger protocols, multi-wave designs, longitudinal collection. Scoped to your study design, priced within the band. | €13 to €17 / hr + 33% |
Every project gets a fixed quote up front: expert hours × your rate in the band, plus the flat 33% service fee, one number, locked before work begins and unchanged through delivery. No other charges: no per-seat fees, no setup fees, no minimums beyond the pack size.
Built to pass your ethics review, not just your budget.
Fair pay you can cite
Evaluators are paid their full hourly rate under fair-pay terms. We provide the compensation and consent documentation your IRB or ethics review needs, and wording you can drop straight into an ethics or broader-impacts statement.
Your data stays yours
You keep full ownership of the data and the study design. We never train on, reuse, or resell what the cohort produces for you. An NDA is available on request.
European by default
Every evaluator is EU-resident and engagements run under European data-protection law. A DPA is available on request, and the setup fits a Horizon Europe or ERC data-management plan.
Anonymity built in
Evaluators are identified by pseudonymous rater IDs, so you can release ratings and agreement statistics, including in double-blind submissions, without exposing anyone.
Why the research rate?
The Scholars program is how we train and credential our expert network, and academic protocols reviewed by real scientists are the best training ground it can get. So research groups get the network at a cost-plus research rate, the experts are paid in full, and published work shows what the network can do. That is the whole trade.
Slots are capped at ten projects per quarter so every study gets a properly managed cohort, reviewed in application order and by fit with the network's domains. The free pilot is always open.
All we ask in return
One line in the acknowledgments of any publication that uses the data:
Plus permission to name your institution among our participating research groups. Nothing else.
What research groups ask first.
Do you train on or reuse our data?
No. The data and the study design are yours. We never train on, reuse, or resell what the cohort produces for you, and an NDA is available on request.
Can I use this in my IRB and ethics statement?
Yes. Evaluators are EU-resident and paid their full hourly rate under fair-pay terms. We provide the compensation, consent, and data-protection documentation your ethics review needs.
Can I bring my own protocol and interface?
Yes. Send your task spec in the exact format you would give an in-house evaluator, and use your own interface or a shared one we set up. The study design stays under your control.
How do you handle quality and rater disagreement?
Every evaluator is trained on your protocol before production work, and items run with rubric-graded QA. On request we run dual independent review and report inter-rater agreement so you can see where raters diverged.
What is the smallest study you take?
The free 5-task pilot has no minimum and no commitment. Paid runs start at the Research Pack of 200 expert hours. Most groups pilot first, then size the full study from the output.
How fast is turnaround?
Pilot data comes back within two weeks. Full studies run on standard two-to-six-week windows, with compressed timelines available at the top of the rate band.
How does billing work?
One fixed quote before work begins, locked through delivery. A single euro invoice on delivered, QA-passed hours, purchase-order friendly. No per-seat fees, setup fees, or hidden minimums.
Ten slots per quarter. The pilot is always open.
Two ways in, both starting with the form:
The free 5-task pilot
Send 5 tasks our way and get the completed data back within two weeks, no cost, no commitment. Most groups start here and size the full study from the output.
A quarterly project slot
A staffed cohort for your full protocol at €13 to €17 per expert hour plus the flat 33% service fee. A short call confirms scope, then you get one fixed quote and a start date.
You're in the review queue
A reply follows from marcus@sovrano.ai within one business day. What happens next:
- Pilot: send your 5 tasks in whatever format your evaluators would get; completed data comes back within two weeks.
- Project slot: a short call to confirm scope, then a single fixed quote and a start date for your cohort.
Slots are reviewed in application order, ten per quarter. The pilot has no cap and no commitment.