LLM Evaluation & RLHF Preference Rating by a Software Engineer

About this Gig
I evaluate model output the way an engineer would — because I am one. Nine years building software, seven years rating AI and search systems, including authoring tasks for AI coding agents on Handshake AI's Project Dynamo. WHAT YOU GET • RLHF preference rating — side-by-side response comparison against your rubric, producing clean preference data with consistent judgement across high volumes. • Red teaming — probing for unsafe, biased, and policy-violating output, with documented findings rather than a pass/fail tick. • Hallucination and factuality checking — verifying claims against sources, catching fabricated citations and invented statistics that read as plausible. • Over-refusal detection — flagging legitimate requests your model declines. Most evaluation counts only harmful output, which optimises toward a model that is safe and useless. • Rubric design review — if two careful raters disagree on your guideline, the problem is usually the guideline. I will tell you where it is ambiguous. • Coding task authoring — original, difficulty-calibrated tasks for agent evaluation, with quality rules and metadata. WHY AN ENGINEER I can read the code your model produces, understand the system it runs in, and explain precisely why an output fails rather than only that it did. That distinction matters most on technical content, where a non-engineer rater cannot tell a subtle bug from working code. HOW I WORK Send your guidelines and a calibration set. I will complete it, report my agreement rate against your gold labels, and flag any rubric ambiguity I hit before starting production work. If the agreement is not where you need it, you have lost one batch rather than a thousand. Available for ongoing work or discrete batches. Mumbai (IST), with overlap across European and US hours.
Requirements
To start, I need: 1. Your evaluation guidelines or rubric — however rough. If you do not have one yet, tell me what "good" means for your use case and I will draft one for your review. 2. A calibration set of 20–50 items with gold labels, so we can both confirm my judgement matches yours before production work begins. 3. Access to whatever tooling you use, or a spreadsheet or JSONL template if you would rather I work in a file. 4. Volume and turnaround — how many items, by when, and whether this is a one-off batch or ongoing. 5. Domain and language scope, and any subject areas to avoid. If you are not sure what you need, send me a sample of your model's output and what is going wrong with it. I will tell you what kind of evaluation would actually answer that question.