Active in Production

2. Product Test PDF Extraction & Multilingual Distribution

A fully automated batch pipeline that scans, downloads, and parses technical product test details from SharePoint PDFs, matches items in OpenSearch, and generates localized translations (IT, EN, ES, FR) of the editorial verdict in parallel.

Processing Time A Few Minutes (Compared to 2 hours manual before)
Automation 100% Automated Scheduled via weekly cron trigger
Frequency 4-5 times / month Automatic weekly batch checks
Operational Impact Human approves output Tested and active in production
Click any node to inspect technical parameters
Weekly Scheduler Cron Trigger Node List SharePoint Folder Scan Node Filter Files Exclude Processed Batch Loop Loop Iterator Node Download PDF Get SharePoint Binary AI PDF Extractor OpenAI Vision Agent Catalog Match OpenSearch Verify Split Route DE Master / Localizer Lang Matrix JS Task Builder AI Localizer Parallel OpenAI Translate Export & Upload CSV Exporter & Sync
Node Inspector

Select a step

Click on any block in the flowchart diagram to view technical parameters, n8n node structures, and example output JSON schemas.

Business Context & Engineering Details

The Business Challenge Solved

Every week, technical test reports are generated in PDF format by the German editorial team. Manually processing these PDFs (reading tables, analyzing technical ratings, collecting pros/cons), finding the matched item in the e-commerce database, and localizing summaries into 4 target markets (Italy, UK, Spain, France) required about 2 hours of manual copying and pasting per execution, 4-5 times a month.

The automated pipeline removes this bottleneck entirely: it runs automatically in the background, parses file structures, matches entities against the catalog, and distributes localized CSV feeds to target SharePoint directories in minutes.

Key Project Highlights:

  • ✔ Zero Human Intervention: Automated trigger scans and processes the incoming queue.
  • ✔ AI PDF Parser: OpenAI processes table grids and editorial text, returning structured JSON schemas.
  • ✔ Catalog Mapping: Automatic matches established through OpenSearch queries.

Technical Highlights & Localization

Batch Processing Loop: Files are managed one-by-one inside a n8n loop block. This isolates potential failures and manages API rate limits during bulk file uploads without crashing the entire run.

Parallel Translation Matrix: Instead of translating sequentially (which causes latency and high run times), a custom JavaScript node splits the translation task into parallel branches. The OpenAI translation nodes execute concurrently for Italian, English, Spanish, and French, accelerating completion speeds.

Reliability: This workflow is fully tested, optimized, and actively running in production, saving hours of manual translation and data entry.