Human-coded benchmark
Established codebooks and adjudicated human codes define the constructs and provide reference data for evaluation.
I am developing AI-assisted behavioral coding methods that build on validated human coding systems to analyze parent–child learning interactions at a scale and consistency that manual coding alone cannot achieve.
Researchers remain central to the process, responsible for interpretation, correction, and validation, preserving the theoretical depth and contextual judgment that human observation of these interactions demands.
Program overview
Fine-grained behavioral coding makes parent–child interaction research possible, but it is time intensive and depends on contextual judgment. This ongoing methods program asks where language models and other machine-coding approaches can assist without turning nuanced behavior into unexamined labels.
The work begins with existing human behavioral coding systems and previously coded data from the Early Mathematics Learning Project, Technology-Mediated Learning project, and Etch-a-Sketch project. Those human judgments provide the basis for training, testing, correcting, and validating machine-generated codes.
Methodological principles
Established codebooks and adjudicated human codes define the constructs and provide reference data for evaluation.
A model proposes candidate labels or helps prioritize cases; its output is treated as a hypothesis, not a final judgment.
Researchers examine disagreement, correct errors, and document reliability before machine-supported codes enter analysis.
Data foundations
EMLP
Structured parent–child mathematics interactions coded in repeated behavioral intervals.
TML
Online mathematics interactions coded at the sentence and minute levels.
EAS
Sentence-level coding of talk during a shared spatial problem-solving task.
Development workflow
Translate established behavioral definitions into testable machine-coding tasks. Use human-coded data as training and evaluation material.
Generate candidate codes for comparison with the human benchmark, treating each output as a proposal rather than a final judgment.
Inspect disagreements to locate ambiguity, missing context, and systematic error. Retain human correction and documentation throughout the pipeline.
Evaluate reliability and validity before using machine-supported codes in substantive research.