Towards a Data Annotation Rubric

Why we need a rubric

When eval­u­at­ing data anno­ta­tion ser­vices, the com­par­i­son points are often veloc­i­ty and price. To com­pete on those terms, some ven­dors take short­cuts that dimin­ish the val­ue of the data, or achieve low­er prices through uneth­i­cal treat­ment of the human beings doing the actu­al anno­ta­tion. Dis­cus­sions of the ven­dor’s process­es and their adher­ence to best prac­tices some­times appear on their web­sites or in con­ver­sa­tion, but it can be dif­fi­cult to objec­tive­ly com­pare providers. Hav­ing a wide­ly used rubric shared by the data com­mu­ni­ty can for­mal­ize these def­i­n­i­tions, remind us of the hid­den lay­ers in data sourc­ing, enable apples-to-apples com­par­i­son, and reduce sur­pris­es when the dataset is delivered.

Fur­ther­more, when a project team has select­ed a ven­dor who isn’t the low­est bid, they often need to jus­ti­fy their choice to a pro­cure­ment team. A score­card that objec­tive­ly mea­sures each ven­dor against a set of stan­dard crieria should make it eas­i­er to explain why spend­ing addi­tion­al mon­ey is appro­pri­ate, in pur­suit of essen­tial high-qual­i­ty data.

Ven­dors will also ben­e­fit from such a rubric, which can help them more effec­tive­ly explain to new clients the val­ue that they pro­vide. It will reduce the pres­sure to com­pete sole­ly on time and mon­ey, allow­ing a rich­er ecosys­tem to thrive with dif­fer­ent providers offer­ing dif­fer­ent trade­offs among price, sched­ule, and quality.

Interpreting the rubric

Each cat­e­go­ry presents a sum­ma­ry of prac­tices we have encoun­tered with var­i­ous ven­dors, assigned to four lev­els. (An emp­ty lev­el in a cat­e­go­ry means we have not yet encoun­tered a prac­tice that we would assign to that level.)

  • Excel­lent: This is a best prac­tice and exceeds expectations.
  • Good: This is ade­quate for most cas­es and is the usu­al expectation.
  • Poor: Below expec­ta­tions; a warn­ing sign that the provider may deliv­er poor-qual­i­ty data.
  • Unac­cept­able: A major defi­cien­cy; even one of these usu­al­ly dis­qual­i­fies a provider.

This rubric, while shared, is not to be mind­less­ly applied across all projects. Your team might adjust the “grade” for some items to empha­size what you con­sid­er impor­tant. We wel­come your feed­back on what should be added or judged dif­fer­ent­ly; please send email to agreene@adobe.com.

Those inter­est­ed in a detailed expla­na­tion of many of these top­ics will find Monarch (2021) useful.

The Rubric

Taxonomy/Ontology/Annotation Guide: Versioning
Excel­lentUses seman­tic ver­sion­ing for the instruc­tions. (See semvar.org)
GoodUses time­stamp or oth­er lin­ear ver­sion num­ber for the instruc­tions.
Main­tains a change log.
PoorLacks a for­mal process for track­ing changes and ensur­ing client agree­ment on changes to the taxonomy/ontology/instructions.
Unac­cept­ableMakes uni­lat­er­al changes; makes it dif­fi­cult to keep client and provider ver­sions “in sync”.
Taxonomy/Ontology/Annotation Guide: Language
Excel­lentMan­u­al­ly trans­lates instruc­tions into the anno­ta­tors’ native lan­guage, when appro­pri­ate. The client is giv­en the oppor­tu­ty to review the translation.
GoodWhen anno­ta­tors are not flu­ent in the lan­guage in which the client has writ­ten guide­lines, ensures that the instruc­tions are easy to under­stand, with client col­lab­o­ra­tion and approval.
PoorDoes not review instruc­tions for read­abil­i­ty by non-flu­ent speak­ers of the lan­guage in which the guide­lines are writ­ten, even when that is needed.
Unac­cept­ableRelies on machine trans­la­tion of instruc­tions with­out man­u­al verification.
Taxonomy/Ontology/Annotation Guide: Questions
Excel­lentHas a UI in the anno­ta­tion tool for anno­ta­tor ques­tions and client respons­es.
Con­sid­ers anno­ta­tor feed­back impor­tant and includes that time in their pay.
GoodCan col­lect anno­ta­tors’ ques­tions in the UI, but respons­es deliv­ered exter­nal­ly.
Pro­vides client with oppor­tu­ni­ty to test-dri­ve the anno­ta­tors’ experience.
PoorUses an out-of-con­text sys­tem (e.g., a shared spreasheet) for anno­ta­tor queries.
Unac­cept­ableLacks any way for anno­ta­tors to ask questions.
Taxonomy/Ontology/Annotation Guide: Test­ing and Refinement
Excel­lentHas experts in anno­ta­tion tech­niques and HCI review the guide­lines and sug­gest improve­ments to increase accu­ra­cy and reduce cog­ni­tive load.
Coach­es under­per­form­ing anno­ta­tors; fir­ing them only as a last resort.
GoodTests the instruc­tions for clar­i­ty and com­plete­ness with small group of anno­ta­tors before scal­ing up.
Anno­ta­tors have option to respond “can’t answer”.
Pro­vides anno­ta­tors suf­fi­cient time to under­stand the task.
Dis­tin­guish­es between train­ing mate­r­i­al and the taxonomy.
PoorInsists that guide­lines must cov­er every even­tu­al­i­ty, even though that adds cog­ni­tive costs while cov­er­age only asymp­tot­i­cal­ly approach­es being complete.
Unac­cept­ableRush­es to scale up anno­ta­tion with­out first col­lect­ing data that con­firms that the guide­lines are clear and are being con­sis­tent­ly applied.
Eth­i­cal treat­ment of anno­ta­tors: Payment
Excel­lentFull-time, paid salary, with benefits.
GoodPart-time, paid hourly, includ­ing time spent in train­ing and on breaks.
PoorPaid per anno­ta­tion, works out to a liv­ing wage in the annotator’s location.
Unac­cept­ablePaid per anno­ta­tion, works out to below liv­ing wage in the annotator’s location.
Eth­i­cal treat­ment of anno­ta­tors: Work conditions
Excel­lentAnno­ta­tion inter­face is acces­si­ble.
Coach­es under­per­form­ing anno­ta­tors, fir­ing them only as a last resort.
GoodPro­vides anno­ta­tors with appro­pri­ate equip­ment (e.g., large mon­i­tors).
Pro­vides breaks and vari­ety to avoid fatigue.
PoorNeglects ergonom­ic needs of anno­ta­tors. Sets unre­al­is­tic quotas.
Unac­cept­ableLacks appro­pri­ate pan­dem­ic safe­ty precautions.
Assess­ing anno­ta­tion quality
Excel­lentProvider’s qual­i­ty team man­u­al­ly reviews a ran­dom sam­ple of data fre­quent­ly.
Pro­vides dash­board (updat­ed dai­ly) for client to mon­i­tor detailed met­rics.
In mul­ti-stage work­flows, anno­ta­tors can flag bad results from pre­vi­ous stages.
Builds mod­els to pre­dict anno­ta­tion errors based on meta­da­ta.
Reports accu­ra­cy esti­mates for each task with confidence/credibility intervals.
GoodIden­ti­fies high-risk items for addi­tion­al man­u­al review.
Sets and meets high­er qual­i­ty goals for test data.
Pro­vides UI for “accep­tance test­ing” by client on a rapid cadence.
Reports accu­ra­cy esti­mates for each task using a sta­tis­ti­cal­ly valid approach.
PoorUses only sim­plis­tic inter-anno­ta­tor agree­ment to mon­i­tor dataset con­sis­ten­cy.
Fails to revis­it ear­li­est anno­ta­tions once anno­ta­tors have gained experience.
Unac­cept­ableFlawed data is dis­card­ed (which can intro­duce bias) instead of being cor­rect­ed.
Task-lev­el accu­ra­cy not report­ed or lacks an expla­na­tion of how it is computed.
Assess­ing anno­ta­tor reliability
Excel­lentUses sta­tis­ti­cal tests to iden­ti­fy out­lier anno­ta­tors for each question.
GoodReg­u­lar­ly adds items whose cor­rect answer is known, to mon­i­tor anno­ta­tor qual­i­ty through­out the project.
Uses sta­tis­ti­cal tests to monitor/identify ques­tions with high dis­agree­ment.
Uses sta­tis­ti­cal tests and track­ing of IP address­es to iden­ti­fy bots or col­lu­sion between sup­pos­ed­ly inde­pen­dent annotators.
PoorUses sim­plis­tic IAA to iden­ti­fy out­lier anno­ta­tors. Rejects minor­i­ty opin­ions out of hand (instead of try­ing to under­stand the cause for the disagreement).
Unac­cept­ableLacks IAA or sta­tis­ti­cal mon­i­tor­ing. Fails to exclude or review data pre­vi­ous­ly obtained from anno­ta­tors who turn out to be unreliable.
Merging/adjudication of indi­vid­ual annotations
Excel­lentMerg­ing strat­e­gy accounts for indi­vid­ual anno­ta­tors’ pre­vi­ous accu­ra­cy. (E.g., reweight­ing of indi­vid­u­als or mod­el­ing cor­re­la­tions among anno­ta­tors.)
Anno­ta­tors can review/revisit their work before it is finalized.
GoodClient can spec­i­fy merg­ing strat­e­gy (e.g., medi­an, or pri­or­i­ty vot­ing, etc.)
Num­ber of anno­ta­tors per task clear­ly spec­i­fied and sta­tis­ti­cal­ly justified.
PoorDeci­sions depend only on major­i­ty vote of annotators.
Unac­cept­ableSome deci­sions are a sin­gle annotator’s opin­ion. (Note: When anno­ta­tor has demon­trat­ed high reli­a­bil­i­ty or is a des­ig­nat­ed SME, this may be defensible.)
Data Deliv­ery
Excel­lentPro­vides detailed raw data includ­ing anno­ta­tor ID, date+time of anno­ta­tion, elapsed time for anno­ta­tion, annotator’s loca­tion and/or locale (for mod­el­ing sources of error such as unan­tic­i­pat­ed cul­tur­al bias or a poor trans­la­tion of the instruc­tions), pre­vi­ous ver­sions for this item from this anno­ta­tor, ver­sion num­ber of the instruc­tions under which each datum was col­lect­ed and anno­tat­ed (plus merged anno­ta­tion data).
GoodPro­vides merged data and indi­vid­ual respons­es but with incom­plete metadata.
PoorPro­vides merged data plus indi­vid­ual respons­es but with min­i­mal metadata.
Unac­cept­ablePro­vides merged data only. Does not exclude or revis­it pre­vi­ous­ly gath­ered data from anno­ta­tors who turn out to be unreliable.
Acquir­ing Unla­beled Data: Ethics
Excel­lentObtains informed con­sent from data providers.
Offers com­pen­sa­tion to con­tent cre­ators when appropriate.
GoodRelies on sources in the pub­lic domain and CC-style licenses.
PoorScrapes the web for pub­licly vis­i­ble data, rely­ing on Fair Use carve-outs in copy­right law in the rel­e­vant geographies.
Unac­cept­ableVio­lates copy­right law in the rel­e­vant geo­gra­phies.
Does not com­ply with pri­va­cy laws such as GDPR.
Acquir­ing Unla­beled Data: Bias and Domain Shift
Excel­lentUses ML approach­es such as clus­ter­ing to mon­i­tor dis­tri­b­u­tion for soci­etal bias and domain mis­match, and resam­ples as needed.
GoodUses heuris­tics to mon­i­tor dis­tri­b­u­tion for soci­etal bias and domain mis­match, and resam­ples as needed.
PoorOnly mon­i­tors for domain shift.
Unac­cept­ableDoes not mon­i­tor for bias or domain shift.
Selec­tion Function/Prioritizing Items
Excel­lentPro­vides meth­ods for (a) dis­tri­b­u­tion sam­pling using provider’s embed­dings, and (b) uncer­tain­ty sam­pling using provider’s off-the-shelf model.
GoodPro­vides API to allow client to pri­or­i­tize data (dis­tri­b­u­tion, uncer­tain­ty based on mod­el under devel­op­ment).
Pro­vides con­trol over ratio of sam­pling meth­ods (ran­dom, dis­tri­b­u­tion, uncer­tain­ty) and allows for that to be updat­ed over time.
PoorDoes not empow­er client to con­trol pri­or­i­ti­za­tion or selec­tion with­in the queue.
Unac­cept­ableUses file­names or data­base IDs to con­trol the order of anno­ta­tion (because these may encode meta­da­ta, such as when items are in chrono­log­i­cal order, and this may prime or bias the annotators).
Start­ing with Seed­ed Data
Excel­lentCom­pares seed­ed and unseed­ed (con­trol-group) anno­ta­tion tasks to mea­sure impact of anchor­ing bias.
Can seed data via in-house com­pu­ta­tion­al approach­es if client desires.
GoodCan seed anno­ta­tions using data pro­vid­ed by client.
Poor-
Unac­cept­ableAuto­mates anno­ta­tion or seeds data via heuris­tics, mod­els, or algo­rithms with­out client’s knowl­edge and assent.

Rubric ver­sion: 0.3.2, last saved 2021-09-30 19:53

Leave a Reply

Your email address will not be published. Required fields are marked *