Supporting Product Discovery ThroughLLM-Based Product Catalog Enrichment andGraph-Based Recommendations : A Case Study of Automotive E-Commerce
Wasim Said · Digitala vetenskapliga arkivet (Diva) (Karlstad University) · 2026
This thesis investigates how LLM-based product-catalog enrichment and graph-based recommendation logic can support product discovery in a resource-constrained automotive e-commerce setting. The case is j-spec.se, a Swedish aftermarket car-parts webshop with a local API snapshot of 39,329 products. After removing 202 rows with empty product names,39,127 products were processed through a pipeline that cleaned catalog text, extracted andnormalized candidate product attributes with LLM support, generated semantic text andembeddings, and ingested the enriched catalog into Neo4j. The resulting graph represented products, brands, categories, specifications, vehicles, engines, variants, and alternative-product relations. The evaluation used offline evidence: source-field preservation checks, validation diagnostics, graph coverage, recommendation candidate coverage, sparse order-derived co-purchase pairs, and selected frontend examples. The pipeline produced structured enrichment for the full cleaned catalog, including specifications for 38,768 products, claims for 34,660products, and extracted fitment for 18,021 products. EAN and weight values were preservedwhere those fields existed in the API data. Graph construction produced 39,127 Productnodes, 293,290 HAS_SPEC relationships, and 28,651 HAS_FITMENT relationships.Complete-your-build candidates were available for 16,757 of 17,957 products with fitmentedges. The results support offline feasibility: enrichment and graph structure can createinspectable recommendation candidates despite uneven catalog data. They do notdemonstrate live customer behavior, production recommendation quality, or business impact.