A large proportion of clinical diagnoses can be derived from language-based interviews alone. These diagnostic interviews are typically conducted by clinicians through doctor-patient interactions during in-person or remote visits. While these interactions are the gold standard for symptom assessment, they can often suffer from financial, geographic, and systemic barriers that limit their accessibility. Current language models (LMs) have demonstrated strong differential diagnosis assessment capabilities when evaluated on curated medical case-studies, highlighting their potential to support the diagnostic process. However, existing evaluations have largely relied on curated, highly detailed and sometimes synthetic patient vignettes, which may not reflect real world experience and clinical presentation variability. These evaluations do not capture how everyday patients report their health symptoms, for example with varying levels of medical literacy, incomplete information, and other complexities that arise through natural conversation. This represents a key gap, leading to uncertainty of how LMs might perform in real-world contexts.To address this gap, we conduct an in-situ comparative research study of a set of experimental conversational prototype AI agents designed to explore how conversational AI might conduct end-to-end symptom interviews and differential diagnostic assessment for research benchmarking purposes. In our recent research paper, “SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment”, we share results from a randomized national scale study (n=13,917) in which consented research participants interact with one of five possible Gemini Flash 2.0 SymptomAI agents. All diagnoses, labels, and disease associations generated during the study were for research analysis only and did not constitute confirmed clinical diagnoses or official medical assessments. Two weeks after their interaction with the AI agents, we asked research participants to report any diagnoses they received from a visit with a healthcare provider. Using this data, we conducted a clinical expert annotation study comparing SymptomAI’s diagnostic performance relative to real clinicians' medical assessments.After assessing the accuracy of SymptomAI’s differential diagnoses (DDx), we further compare SymptomAI’s diagnoses against biosignals from participants’ Fitbit wearable devices in the time leading up to their conversation with SymptomAI. We show that SymptomAI conversations that led to diagnosis with an infectious disease etiology coincide with physiological trends that may indicate an immune response, suggesting further evidence of SymptomAI’s performance.