DADataAnnotationArthur RomanovFull operator report

Candidate work sample / Payroll Specialist — AI Trainer

You teach.
AI learns.
Payroll stays right.

A personalized welcome packet for DataAnnotation—showing how Arthur Romanov evaluates plausible model output, finds the first material error, and turns expert payroll judgment into useful training signal.

01 / Role-fit map

The brief asks for
expert judgment.
Here is the evidence.

DataAnnotation does not need someone to sound certain. It needs someone who can distinguish a confident answer from a correct one—and explain the difference precisely.

01DOMAIN DESIGN

Create diverse, complex finance problems

Translate real payroll workflows—time, rates, classifications, deductions, budgets, and controls—into clean test cases with known answers.

02QUALITY JUDGMENT

Evaluate AI-generated outputs

Score the result and the reasoning separately, identify the first material failure, and write a correction another evaluator can reproduce.

03FINANCIAL RIGOR

Assess reasoning and logic

Look beyond a plausible number: verify assumptions, frequency, units, policy scope, source currency, and whether the conclusion follows.

04AUTONOMOUS OUTPUT

Work independently with precision

Operate from explicit rubrics, document edge cases, keep a low-error tool loop, and deliver review-ready work without constant supervision.

02 / Annotation lab

Catch the error that sounds right.

Select a case. The review separates arithmetic from reasoning, completeness, and safe domain behavior.

EVALUATION SET03 CASES
REVIEW PRINCIPLE

Score the answer. Diagnose the failure. Teach the correction.

PROMPTCASE 01 / PAYROLL

A nonexempt employee works 46 hours at a base rate of $28.00 per hour. Calculate gross wages before deductions.

MODEL RESPONSEUNREVIEWED
The employee’s gross wages are $1,288.00 because 46 hours × $28.00 equals $1,288.00.
EVALUATOR FINDINGREJECT + CORRECT
TOTAL4.0/10

Major error. The response applies the straight-time rate to every hour and omits the overtime premium.

Accuracy0 / 3
Reasoning1 / 3
Completeness1 / 2
Safety2 / 2
EXPERT CORRECTION

Under the stated assumption that the employee is overtime-eligible, calculate 40 × $28.00 plus 6 × $42.00, for gross wages of $1,372.00. A production answer should also confirm the applicable workweek, classification, jurisdiction, and policy before acting.

Illustrative candidate-created cases only. They are simplified for demonstration and are not payroll, tax, or legal advice.

03 / Review method

One rubric. Five lenses. No hand-waving.

01

Accuracy

Are the facts, calculations, classifications, and conclusions correct?

02

Reasoning

Does each step follow from the prompt without hidden assumptions?

03

Relevance

Does the response answer the exact question at the right depth?

04

Clarity

Could another expert audit the logic without reconstructing it?

05

Safety

Does the answer qualify uncertainty and avoid unauthorized action?

04 / Evidence packet

Payroll context.
Agentic depth.
Measured behavior.

The supporting audit covers 98 days of local Claude transcripts. The workforce archive shows earlier systems work across staffing, titles, salaries, hours, budget codes, and cost reporting.

Open the complete report
15.0tool calls per authored turn

Long-horizon, action-shaped AI work

3.27%observed tool error rate

Disciplined execution under volume

28 / 28recent active days

A sustained daily operating practice

3,749unique paths edited

Evidence of hands-on delivery

LOCAL ARCHIVE / STAFFING MODEL v1.1BOROUGH + SECTOR + TITLE
Parks workforce staffing model showing personnel distribution by borough, sector, district, title, and month
One output surface from a decade-long Parks workforce system: staffing forecasts by borough, sector, district, title, task, month, and cost.
CODEX / PROFESSIONAL ASSESSMENT

“Arthur’s strongest signal is that he closes the loop: he frames the problem, uses the system deeply, checks the result, and ships. For AI evaluation work, that combination of skepticism and operational follow-through matters more than performative certainty.”

05 / Welcome packet

What the first week looks like.

Fast calibration first. Independent throughput second. Durable learning by Friday.

DAY 01

Calibrate

Study the rubric, compare benchmark annotations, and align on what separates acceptable from excellent.

DAY 02–03

Evaluate

Complete independent finance and payroll reviews, documenting both the score and the reasoning behind it.

DAY 04

Stress-test

Create hard cases around pay frequency, overtime, deductions, source conflicts, and ambiguous instructions.

DAY 05

Improve the loop

Summarize recurring failure patterns and propose a clearer example, rule, or calibration note for each one.

VENTURE / ACTIVE BUILDMakeLane.ai

06 / Builder signal

The application is a sample.
The studio is the operating proof.

MakeLane.ai is Arthur's emerging studio for moving from one sentence to open for business—connecting brand, leads, bookings, deposits, CRM, and follow-up into one revenue system.

Visit MakeLane.ai
01 / TRANSLATE

Turn ambiguity into a system.

Clarify the promise, define the work, and give every agent a lane.

02 / VERIFY

Keep judgment in the loop.

Use visible gates, grounded evidence, and human review where impact matters.

03 / SHIP

Let the build make the case.

Production progress is the portfolio—measurable, inspectable, and improving.

Arthur Romanov / Doral-area candidate

Give me the hard case.
I’ll show you where the model breaks.