In Development (WIP)

1. Web Review Scraping & Catalog Match Engine

An intelligent automation designed to discover new web reviews, extract structured metadata using hybrid scraping (JSON-LD / AI Fallback), execute fuzzy catalog queries, and semantically validate matches via AI agents.

Scraping Time ~5 Minutes (Estimated 1 day manual before)
Automated Translation -2 Hours / Run (In development phase)
Frequency 5-6 times / month Scheduled & event-driven
Testing Status Active Debugging Not yet tested in prod
Click any node to inspect technical parameters
Start Execution Scheduled Trigger Discover URLs Sitemap RSS Scanner Review Crawler HTML Page Downloader Strategy Router JS Code Selection JSON-LD Parser Metadata HTML Extract HTML Scraper CSS Scraper Fallback AI Extraction OpenAI HTML Reader Data Normalizer JS Schema Normalizer Catalog Lookup OpenSearch Fuzzy Query AI Match Verify OpenAI Semantic Decision AI Fazit & Translate Work in Progress (WIP) SharePoint Upload CSV Feed Exporter
Node Inspector

Select a step

Click on any block in the flowchart diagram to view technical parameters, n8n node structures, and example output JSON schemas.

Business Context & Engineering Details

The Business Challenge Solved

Traditionally, tracking and mapping reviews from external blogs and portals required marketing teams to manually scroll pages, copy text, and format results. This manual process took about 1 full working day per portal.

Traditional scraping templates frequently break when website layouts change. Furthermore, associating a review with the correct catalog item is difficult due to name variations (e.g., 'Aero 45 Black' vs 'Aero 45L Trekking Backpack').

Key Project Highlights:

  • ✔ Hybrid Scraping: Reliable JSON-LD parsing with automated CSS/AI fallbacks.
  • ✔ Zero False Matches: Semantic AI validation blocks incorrect catalog mappings.
  • ✔ Scalability: Instant run in 5 minutes instead of a full manual working day.

Architectural Decisions & Costs

API Cost Optimization: Calling OpenAI LLMs is costly. The workflow is optimized to invoke the AI only as a fallback (when JSON-LD is missing) and for final semantic confirmation. Supplying clean, structured data to the model minimizes context lengths, reducing average execution costs to less than a cent per run.

OpenSearch Query Optimization: OpenSearch fuzzy searches filter down 30,000 catalog entries to a top-3 candidate list, meaning the OpenAI agent only has to evaluate a highly refined list, eliminating hallucinations and latency.

WIP Status: Currently, the translation module is in development on n8n. Once integrated, it will automate translation and formatting, saving 2 hours of manual translation per run (runs 1-2 times per month).