Project Activities
The team will create
- a longitudinal dataset on schools’ racial/ethnic and economic composition since 1991 (since 1967 for some schools);
- a longitudinal dataset on racial/ethnic and economic segregation between schools within school districts, counties, metropolitan areas, commuting zones, states, and the US (and, for larger geographies, segregation between school districts);
- technical documentation describing the methods used to create the datasets; and
- learning resources including data explainers, statistical code, methodological briefs, and training webinars. The investigators will publish data products, associated documentation, and resources to aid in the use and interpretation of the data in a publicly accessible data archive, linked via a public-facing website.
Structured Abstract
Research design and methods
To address our research aims, we employ several methods. First, we will use statistical and/or machine learning techniques for anomaly detection to identify erroneous school demographic data in the Common Core of Data (CCD). Second, we will develop longitudinal imputation models appropriate for panel data to impute and replace missing and erroneous data, and we will validate the results of our imputation using secondary data analyses of other NCES surveys. We will also use probabilistic matching techniques to match schools in CCD and Office of Civil Rights (OCR) data, producing the longest-term school demographic dataset available. Third, we will estimate multiple measures of segregation between schools and school districts at a range of administrative and geographic levels, accounting for uncertainty in analyzing sample and imputed demographic data. Fourth, we will incorporate user-testing into our methods to iteratively improve on our modeling and data products to ensure usability. Fifth, we will develop iterative workflows to keep the data products updated, including techniques to update our anomaly detection and imputation models with additional annual data releases.
User Testing: We will convene a group of education researchers who use school demographic data and/or segregation data to participate in testing sessions. We will recruit this group through personal invitation to collaborators and colleagues, recruitment via professional organization listservs, and social media. We will share preliminary datasets, technical documentation, and user resources. We will solicit feedback through several mechanisms. First, we will ask users to complete simple data tasks and share their log files and output. Second, we will encourage users to attempt to analyze the data in a way that would be typical of their research needs and respond to a brief survey via Google forms adapted from the System Usability Scale (SUS) (Lewis 2018). Finally, we will invite users to provide us with open-ended feedback on our methodology, data products, and user resources via email or a brief Zoom interview. We will incorporate user feedback into revisions of both the data, accounting for any errors or suggestions our users find, and in the user resources to improve usability.
People and institutions involved
IES program contact(s)
Project contributors
Products and publications
We will publish data products, technical documentation describing our methodology and procedures, and resources to facilitate use of the data in education research via the Segregation Explorer website (http://edopportunity.org/segregation). The website will link to the Stanford Digital Repository, where the data will be stored in a publicly-accessible, secure, and institutionally-maintained data archive. We plan to disseminate the products in several ways. First, we will work with the Stanford and UCLA communications offices to develop communications plans to advertise new data releases. We will also share updates via social media and email announcements to data users, whose emails we will collect when they access our data. Second, to engage education researchers mainly employed in academic settings, we will submit session proposals to hold data workshops at relevant professional organizations’ annual conferences (e.g., American Educational Research Association (AERA), Society for Research on Educational Effectiveness (SREE), and the American Sociological Association (ASA)). Third, to reach users employed in other settings, we will leverage our connections with social change organizations and practitioners to conduct workshops for education researchers at district and state educational agencies, advocates, and others. These workshops will likely be geared more toward using the end products than the methodologies used to create them, given what we know about this audience’s interests. We will also hold virtual workshops, advertised to education researchers through our social networks, social media, and professional organization listservs. These workshops will be recorded and posted on our website for all users to access.
Use in Applied Education Research: School demographic composition and school segregation are topics of frequent study in education research. For example, scholars of student assignment policies, educational inequality, peer effects, and school inputs and outputs frequently consider these variables. Moreover, even scholars not centrally interested in these concepts frequently use school demographic composition or segregation as control variables in their analyses. Therefore, producing comprehensive, accurate, longitudinal datasets on school demographic composition and school segregation will improve education research beyond typical practices of using the CCD without attending to quality issues or estimating school segregation without appropriate expertise.
Project website:
Related projects
Questions about this project?
To answer additional questions about this project or provide feedback, please contact the program officer.