Lab or Field? AMethod-Comparison Study of Lighthouse and CrUXMeasurements on the Turkish Web
1st.com.tr Araştırma Birimi · Zenodo (CERN European Organization for Nuclear Research) · 2026
To determine how well laboratory (Lighthouse, throttled) performance measurements predict real-userfield (Chrome UX Report, CrUX p75) experience. Method: For 7,403 Turkish sites with both laboratory and fielddata, Largest Contentful Paint (LCP), First Contentful Paint (FCP) and Cumulative Layout Shift (CLS) werematched; agreement was assessed with Pearson and Spearman correlations, Bland-Altman analysis, and Cohen'skappa over the Core Web Vitals 'good/needs-improvement/poor' classification. 1st.com.tr Results: The laboratorysystematically and substantially over-reported loading metrics (LCP: lab median 10,876 ms vs field 2,247 ms). Thelab–field relationship for LCP was weak (ρ = 0.26) and classification agreement did not exceed chance (κ ≈ 0.00);85.6% of sites were classified worse in the lab than in the field, and while only 3.1% of sites were 'good' in the lab,59.6% were 'good' in the field. FCP showed similarly weak agreement (κ = 0.01); only CLS reached moderateagreement (κ = 0.33). The overall Lighthouse score correlated only weakly with field LCP (r = −0.25). Conclusion:On the Turkish web, default laboratory Lighthouse loading metrics are a poor proxy for real-user experience;performance decisions should rest on field data, with the laboratory reserved for diagnosis.