Create diverse, complex finance problems
Translate real payroll workflows—time, rates, classifications, deductions, budgets, and controls—into clean test cases with known answers.
Candidate work sample / Payroll Specialist — AI Trainer
A personalized welcome packet for DataAnnotation—showing how Arthur Romanov evaluates plausible model output, finds the first material error, and turns expert payroll judgment into useful training signal.
01 / Role-fit map
DataAnnotation does not need someone to sound certain. It needs someone who can distinguish a confident answer from a correct one—and explain the difference precisely.
Translate real payroll workflows—time, rates, classifications, deductions, budgets, and controls—into clean test cases with known answers.
Score the result and the reasoning separately, identify the first material failure, and write a correction another evaluator can reproduce.
Look beyond a plausible number: verify assumptions, frequency, units, policy scope, source currency, and whether the conclusion follows.
Operate from explicit rubrics, document edge cases, keep a low-error tool loop, and deliver review-ready work without constant supervision.
02 / Annotation lab
Select a case. The review separates arithmetic from reasoning, completeness, and safe domain behavior.
Score the answer. Diagnose the failure. Teach the correction.
A nonexempt employee works 46 hours at a base rate of $28.00 per hour. Calculate gross wages before deductions.
The employee’s gross wages are $1,288.00 because 46 hours × $28.00 equals $1,288.00.
Major error. The response applies the straight-time rate to every hour and omits the overtime premium.
Under the stated assumption that the employee is overtime-eligible, calculate 40 × $28.00 plus 6 × $42.00, for gross wages of $1,372.00. A production answer should also confirm the applicable workweek, classification, jurisdiction, and policy before acting.
Illustrative candidate-created cases only. They are simplified for demonstration and are not payroll, tax, or legal advice.
03 / Review method
Are the facts, calculations, classifications, and conclusions correct?
+Does each step follow from the prompt without hidden assumptions?
+Does the response answer the exact question at the right depth?
+Could another expert audit the logic without reconstructing it?
+Does the answer qualify uncertainty and avoid unauthorized action?
+04 / Evidence packet
The supporting audit covers 98 days of local Claude transcripts. The workforce archive shows earlier systems work across staffing, titles, salaries, hours, budget codes, and cost reporting.
Open the complete reportLong-horizon, action-shaped AI work
Disciplined execution under volume
A sustained daily operating practice
Evidence of hands-on delivery

CODEX / PROFESSIONAL ASSESSMENT“Arthur’s strongest signal is that he closes the loop: he frames the problem, uses the system deeply, checks the result, and ships. For AI evaluation work, that combination of skepticism and operational follow-through matters more than performative certainty.”
05 / Welcome packet
Fast calibration first. Independent throughput second. Durable learning by Friday.
Study the rubric, compare benchmark annotations, and align on what separates acceptable from excellent.
Complete independent finance and payroll reviews, documenting both the score and the reasoning behind it.
Create hard cases around pay frequency, overtime, deductions, source conflicts, and ambiguous instructions.
Summarize recurring failure patterns and propose a clearer example, rule, or calibration note for each one.
06 / Builder signal
MakeLane.ai is Arthur's emerging studio for moving from one sentence to open for business—connecting brand, leads, bookings, deposits, CRM, and follow-up into one revenue system.
Visit MakeLane.aiClarify the promise, define the work, and give every agent a lane.
Use visible gates, grounded evidence, and human review where impact matters.
Production progress is the portfolio—measurable, inspectable, and improving.
Arthur Romanov / Doral-area candidate