Sorting Algorithms for Qualitative Data to Recover Latent Dimensions with Crowdsourced Judgments: Measuring State Policies for Welfare Eligibility under TANF.

James Honaker, Michael B. Berkman, Chris Ojeda, Eric Plutzer · 2013

The Quicksort and Bubble Sort algorithms are commonly implemented procedures in computer science for sorting a set of numbers from low to high in an eEcient number of processes using only pairwise comparisons. Becauseofsuchalgorithms’relianceonpairwisecomparison,theylendthemselvestoanyimplementation where a simple judgment requires selecting a winner. We show how such algorithms, adapted for stochastic measurements, are an eEcient way to harness human ”crowdsourced” coders who are willing to make brief judgmentscomparingtwopiecesofqualitativeinformation(here,sentencesoftext)touncovertheunderlying dimension or structure of the qualitative sources. Asademonstrationoftheabilityofourapproach,weshowthatcorrectlystructurednon-expertjudgments ofthelevelofdemocratizationincountriesrecoversthesameinformationthatalternateexpertscalesofdemocratization ‐with large cost and time‐ estimate. Our key motivating implementation involves a large collection of written policies describing conditions for eligibility for welfare (TANF) in each state in the US. An open question in the literature on welfare is the existenceofarace-to-the-bottom,whichnecessitatesmeasuringthegenerosityofcomplexsetsofeligibilityrules that diAer from state to state and across time. Existing approaches in the literature have attempted to scale or rank state generosity in welfare policies by either constructing large coding questionnaires (Fellowes and Rowe, 2004) and summing responses, or by attempting factor analysis of all possible raw data on welfare policy (De Jong et al, 2006). The Arst approach requires understanding all the dimensions that are relevant before constructing the survey implement. The second requires converting all policy documents and rules into quantitative measures. Weshowhowtoobtainarankingofstatewelfaregenerositywithoutdoingharmtothequalitativenatureof the sources, and without leveraging expert knowledge to sort the vast collection of textual sources. We present ”crowdsourced” human coders sentences describing one policy measure in each of two states, and ask them which of the two is the more generous (or more Cexible, or more lenient) welfare rule. The optimal set of pairwisecomparisonsiscontinuouslychosenbythesortingalgorithm. Wecomparetherankingsofstatepolicies

Read the paper · More papers on PaperTik