by Khairul Muhtadin
Transform raw investment memorandums and financial decks into comprehensive, professional Due Diligence (DD) PDF reports. This workflow automates document parsing via LlamaParse, enriches internal data with real-time web intelligence using Decodo, and utilizes an AI Agent to synthesize structured financial analysis, risk assessments, and investment theses. Why Use This Workflow? Time Savings:** Reduces initial deal screening and report generation from 6–8 hours of manual analysis to under 5 minutes. Accuracy & Depth:** Employs a multi-query RAG (Retrieval-Augmented Generation) strategy that cross-references internal deal documents with verified external web evidence. Cost Reduction:** Eliminates the need for expensive junior analyst hours for preliminary data gathering and document summarization. Scalability:** Effortlessly processes multiple deals simultaneously, maintaining a consistent reporting standard across your entire pipeline. Ideal For Venture Capital & Private Equity:** Rapidly assessing incoming pitch decks and CIMs (Confidential Information Memorandums). M&A Advisory Teams:** Automating the creation of standardized target company profiles and risk summaries. Investment Analysts:** Generating structured data from unstructured PDFs to feed into internal valuation models. How It Works Trigger: A webhook receives document uploads (PDF, DOCX, PPTX) via a custom portal or API. Data Collection: LlamaParse converts complex document layouts into clean Markdown, preserving tables and financial structures. Processing: The workflow generates a unique "Deal ID" based on filenames to ensure data isolation and implements a caching layer via Pinecone to avoid redundant parsing. Intelligence Layer: Web Enrichment: The workflow derives the target company name and uses Decodo to scrape official websites for "About" and "Commercial Risk" data. Multi-Query RAG: An OpenAI-powered agent executes six specific retrieval queries (Financials, Risks, Business Model, etc.) to gather evidence from all sources. Output & Delivery: Analysis is mapped to a structured template, rendered into a professional HTML report, and converted to a high-quality PDF using Puppeteer. Storage & Logging: The final report is uploaded to Cloudflare R2, and a public, secure URL is returned to the user instantly. Setup Guide Prerequisites | Requirement | Type | Purpose | | --- | --- | --- | | n8n instance | Essential | Core automation and workflow orchestration | | LlamaIndex Cloud | Essential | High-accuracy document parsing (LlamaParse) | | Pinecone | Essential | Vector database for document and web evidence storage | | OpenAI API | Essential | LLM for embeddings and expert analysis (Embedding Small & GPT-5.2) | | Decodo API | Essential | Real-time web searching and markdown scraping | | R2 Bucket | Essential | Secure storage for the generated PDF reports | Installation Steps Import the JSON file to your n8n instance. Configure credentials: OpenAI: Add your API key for embeddings and the Chat Model. Pinecone: Enter your API Key and Index name (default: poc). LlamaIndex: Add your API key under Header Auth (Authorization: Bearer YOUR_KEY). Decodo: Set up your Decodo API credentials for web search and scraping. AWS S3: Configure your bucket name and access keys. Update environment-specific values: In the "Build Public Report URL" node, update the baseUrl to match your S3 bucket's public endpoint or CDN. Test execution: Send a POST request to the webhook URL with a binary file (e.g., a Pitch Deck) to verify the end-to-end generation. Technical Details Core Nodes | Node | Purpose | Key Configuration | | --- | --- | --- | | LlamaParse (HTTP) | Document Conversion | Uses the /parsing/upload and /job/result endpoints for high-fidelity markdown | | Pinecone Vector Store | Context Storage | Implements namespace-based isolation using the unique dealId | | Decodo Search/Scrape | Web Intelligence | Dynamically identifies the official domain and extracts corporate metadata | | AI Agent | Strategic Analysis | Configured with a "Senior Investment Analyst" system prompt and 6-step retrieval logic | | Puppeteer | PDF Generation | Renders the styled HTML report into a print-ready A4 PDF | Workflow Logic The workflow uses a Multi-Query Retrieval strategy. Instead of asking one generic question, the AI Agent is forced to perform six distinct searches against the vector database (Revenue History, Key Risks, etc.). This ensures that even if a document is 100 pages long, the AI doesn't "miss" critical financial tables or risk disclosures buried in the text. Customization Options Basic Adjustments Report Styling:** Edit the "Render DD Report HTML" node to match your firm's branding (logo, colors, fonts). Analysis Scope:** Modify the AI Agent's prompt to include specific metrics (e.g., "ESG Score" or "Technical Debt Assessment"). Advanced Enhancements Slack/Email Integration:** Instead of just an S3 link, have n8n send the PDF directly to a #new-deals Slack channel. CRM Sync:** Automatically create a new record in HubSpot or Salesforce with the structured JSON output attached. Troubleshooting | Problem | Cause | Solution | | --- | --- | --- | | Parsing Timeout | File is too large for synchronous processing | Increase the "Wait" node duration or check LlamaParse job limits | | Low Analysis Quality | Insufficient context in documents | Ensure documents are text-based PDFs (not scans) or enable OCR in LlamaParse | | PDF Layout Broken | CSS incompatibility in Puppeteer | Simplify CSS in the HTML node; avoid complex Flexbox/Grid if Puppeteer version is older | Use Case Examples Scenario 1: Venture Capital Deal Screening Challenge: A VC associate receives 20 pitch decks a day and spends hours manually summarizing company profiles. Solution: This workflow parses the deck and web-scrapes the startup's site to verify claims. Result: The associate receives a 3-page PDF summary for every deck, allowing them to reject or move forward in seconds. Scenario 2: Private Equity Due Diligence Challenge: Analyzing a 150-page CIM (Information Memorandum) for specific financial "red flags." Solution: The AI Agent is programmed to specifically hunt for customer concentration and margin fluctuations. Result: Consistent risk identification across all deals, regardless of which analyst is assigned to the project. Created by: Khmuhtadin Category: Business Intelligence | Tags: Decodo, AI, RAG, Due Diligence, LlamaIndex, Pinecone Need custom workflows? Contact us Connect with the creator: Portfolio • Store • LinkedIn • Medium • Threads
by Davide
This workflow automatically generates an llms.txt file (following the llmstxt.org specification) for any given website. It uses ScrapegraphAI to crawl and scrape pages, an OpenAI chat model to process content, and finally uploads the generated file via FTP. Key Advantages 1. ✅ Automated llms.txt Generation The workflow fully automates the creation of a compliant llms.txt file, eliminating the need for manual documentation and reducing maintenance time. 2. ✅ AI-Powered Website Understanding Using OpenAI and ScrapeGraphAI, the system intelligently analyzes: Website structure Internal pages Titles and descriptions Content relevance Logical page categorization This produces a high-quality output specifically optimized for AI systems and LLM indexing. 3. ✅ Dynamic Internal Link Discovery The crawler automatically extracts all internal links from the website, making the workflow scalable for: Small business websites Large corporate websites Ecommerce stores Blogs and documentation portals 4. ✅ Intelligent Content Categorization Pages are automatically grouped into meaningful sections such as: Main Pages Services Products Portfolio Blog Company Contact Legal / Optional pages This improves readability and machine interpretability. 5. ✅ Multilingual Support The workflow preserves the original language of the website content, ensuring consistency and localization for international projects. 6. ✅Fully Automated Publishing After generation, the workflow converts the output into a .txt file and uploads it directly to an FTP server or CDN, enabling instant deployment without manual intervention. 7. ✅ Reduced Manual Work* The entire process — from crawling to publishing — is automated inside n8n, significantly reducing operational effort for SEO teams, developers, and AI optimization workflows. 8. ✅ AI & SEO Optimization The generated llms.txt file helps: AI crawlers better understand the website Improve AI discoverability Structure content for LLM consumption Support future AI search indexing strategies 9. ✅ Modular and Scalable Architecture The workflow is built with reusable components: Crawler module Status monitoring AI analysis agent Scraper tool Binary conversion FTP deployment This makes it easy to extend, customize, or integrate into larger automation systems. Ideal Use Cases AI-ready website optimization Automated SEO infrastructure LLM indexing preparation Agency website automation Large-scale multi-site management Documentation platforms AI search visibility enhancement How it works The process begins when the workflow is manually triggered. It then: Starts a crawl of the specified domain using ScrapegraphAI’s smartcrawler. The crawler extracts all internal links from the domain (acting like a sitemap generator). Waits for the crawl to complete (configurable wait time, default 20 units). Checks the crawler’s status – if the crawl is still processing, the workflow waits again; if successful, it proceeds. Extracts the discovered internal links and passes them to an AI agent. Uses an AI agent (with OpenAI GPT) that: Receives the list of internal URLs. Uses a Scraper tool (via ScrapegraphAI) to scrape each URL’s content. Follows a strict prompt to: Analyze the homepage (title, description, language). Extract concise descriptions for each internal page. Group pages into logical sections (Main pages, Services, Portfolio, Contact, Optional, etc.). Generate a clean Markdown file (llms.txt) following the official spec. Converts the Markdown output into a binary file (llms.txt). Uploads the file to an FTP server (configured for BunnyCDN or any FTP storage). Ends the workflow once the upload is complete. The AI agent is explicitly forbidden from inventing content – it must call the Scraper tool for every URL before describing it. The output is pure Markdown, starting with #. Setup steps To use this workflow in n8n, follow these steps: 1. Prerequisites An n8n instance (self-hosted or cloud). A ScrapegraphAI account with API access. An OpenAI account with API key (model used: gpt-5.4-mini – note: this may be a custom/typo; usual models are gpt-4o-mini or gpt-4). An FTP server (traditional FTP, or SFTP if modified). 2. Configure credentials in n8n Go to Credentials in n8n and add: ScrapegraphAI API** Name: ScrapegraphAI account API Key: your ScrapegraphAI API key OpenAI API** Name: OpenAi account (Eure) API Key: your OpenAI API key FTP** Name: FTP BunnyCDN Host, Port, Username, Password (or SSH key) for your FTP server 3. Modify the domain In the Set domain node, change the your_domain to your target domain (e.g., example.com). Do not include https:// – only the domain name. 4. Adjust wait time (optional) In the Wait node, change the amount (default 20) to a higher value if the target site is large or slow to crawl. 5. Update FTP upload path In the Upload to FTP node, update the path field. Currently it is: =/YOUR_PATH/{{$binary.data.fileName}} Change YOUR_PATH to the actual remote directory (e.g., /public_html/). The file will be saved as llms.txt. 6. (Optional) Modify the AI prompt The prompt inside the LLMS.txt Agent node can be adapted for: Different section names Different output structure Different languages Exclusion of certain URL patterns 7. Activate and execute Save the workflow. Toggle Active to enable manual execution. Click ‘Execute workflow’ on the Manual Trigger node. Monitor execution – the workflow will wait for the crawl, then process all pages, and upload the final file. 8. Verify Check your FTP server for the generated llms.txt. Test it by opening in a text editor – it should be pure Markdown starting with # Site name. 👉 Subscribe to my new YouTube channel. Here I’ll share videos and Shorts with practical tutorials and FREE templates for n8n. Need help customizing? Contact me for consulting and support or add me on Linkedin.
by Emir Belkahia
This workflow helps Customer Success Managers and customer success professionals quickly gather intelligence on clients or prospects by analyzing their recent LinkedIn activity via a simple Slack command. Who's it for CSMs, Account Managers, and Sales professionals who need fast, structured insights about a person's LinkedIn presence before a call, meeting, or outreach. What it does (and doesn't do) ✅ It DOES: Fetch recent LinkedIn posts from any profile Analyze posting frequency and cadence patterns Identify top themes and focus areas Extract recent highlights with context Generate a clean HTML report sent via email ❌ It DOESN'T: Access private/non-public LinkedIn content Provide real-time updates (it's a snapshot) Replace actual researches when needed Think of it as: Your personal LinkedIn research assistant that turns a name into actionable intelligence in under a minute. How it works Slack command - Type /check-linkedin [Full Name] in Slack Name validation - AI verifies you provided a full name (not just "John") Profile discovery - Finds the correct LinkedIn profile via Apify Content scraping - Pulls their recent posts (last 20) AI analysis - GPT-4.1 analyzes posting patterns, topics, and highlights Report generation - Creates a formatted HTML email report Email delivery - Sends the intelligence brief to your inbox Set up steps Setup time: ~15 minutes Create or use your existing Slack app and add a Slash Command (it can be done here https://api.slack.com/apps) Configure the webhook URL in your Slack app Connect credentials: Slack OAuth Apify API OpenAI API Gmail OAuth Update the email recipient in "Send report via Email" node Test with a known LinkedIn profile Requirements Slack workspace (with app installation permissions) Apify account with credits OpenAI API key (GPT-4.1 access) Gmail account Apify actors: LinkedIn Profile Finder LinkedIn Post Scraper Cost estimation ~$0.05-0.09 per profile check. You could research 11-20 people for $1. ⚠️ Cost Disclaimer: The costs displayed above are indicative only and may vary significantly depending on which n8n actors you select. Some actors incur monthly charges—for example, one of the two actors used in this workflow costs $35/month. So, I recommend using this actor only when there's a clear business need for it. For cost optimization, consider switching to alternative actors that can deliver similar / simpler functionality at a lower cost. If you plan to use this workflow extensively, I strongly suggest performing a budget assessment and evaluating other actor options to maximize cost efficiency. The workflow uses GPT-4.1-mini for lightweight classification and GPT-4.1 for the heavy analysis to balance quality and cost. Known Limitations Common names have limited accuracy: Very common names (e.g., "John Smith") often fail to identify the correct person accurately. An improved version could support company name in the slash command as an additional input to help narrow down results and improve first-try matching accuracy. 💡 Pro tips Check before important meetings: Run this 15-30 minutes before a call. The email report gives you conversation starters and context about what they care about. Batch your research: If you have multiple clients or prospects, queue them up. Just remember each lookup costs ~$0.05-0.09. Watch your Apify credits: The LinkedIn scrapers are the main cost driver. Monitor your Apify usage if you're doing high volume. Don't spam the same profile: LinkedIn may rate-limit. Space out repeat checks on the same person by at least a few hours. Review the "Quick Scan" section first: The email report starts with key stats and top focus areas. Perfect for a 30-second pre-call prep. What to do after the workflow runs Check your email - Report arrives in 30-90 seconds Review the report - Latest post date, cadence, and top themes Read Recent Activity Summary - High-level overview of their content Dive into Detailed Analysis - Two main topics with keywords and rationale Use it strategically: Reference their recent posts in your outreach Ask about topics they're clearly passionate about Tailor your pitch to their demonstrated interests Avoid generic "saw you on LinkedIn" messages Questions or Feedback? 📧 emir.belkahia@gmail.com 💼 linkedin.com/in/emirbelkahia
by Pawan
This template sets up a scheduled automation that scrapes the latest news from The Hindu website, uses a Google Gemini AI Agent to filter and analyze the content for relevance to the Competitive Exams like UPSC Civil Services Examination (CSE) syllabus, and compiles a structured daily digest directly into a Google Sheet. It saves hours of manual reading and note-taking by providing concise summaries, subject categorization, and explicit UPSC importance notes. Who’s it for This workflow is essential for: UPSC/CSE Aspirants who require a curated, focused, and systematic daily current affairs digest. Coaching Institutes aiming to instantly generate structured, high-quality study material for their students. Educators and Content Creators focused on Governance, Economy, International Relations, and Science & Technology. How it works / What it does This workflow runs automatically every morning (scheduled for 7 AM by default) to generate a ready-to-study current affairs document. Scraping: The Schedule Trigger fires an HTTP Request to fetch the latest news links from The Hindu's front page. Data Curation: The HTML and Code in JavaScript nodes work together to extract and pair every article URL with its title. Content Retrieval: For each identified link, a second HTTP Request node fetches the entire article body. AI Analysis and Filtering: The AI Agent uses a detailed prompt and the Google Gemini Chat Model to perform two critical tasks: Filter: It filters out all irrelevant articles (e.g., sports results, local crime) to keep only the 5-6 most important UPSC-relevant pieces (Polity, Economy, IR, etc.). Analyze: For the selected articles, it generates a Brief Summary, identifies the Main Subject, and clearly articulates Why it is Important for the UPSC Exam. Storage: The AI Agent calls the integrated Google Sheets Tool to automatically append the structured, analyzed data into your designated Google Sheet, creating your daily ready-made notes. Requirements To deploy this workflow, you need: n8n Account: (Cloud or self-hosted). Google Gemini API Key: For connecting the Google Gemini Chat Model and powering the AI Agent. Google Sheets Credentials: For reading/writing the final compiled digest. Target Google Sheet: A spreadsheet with the following columns: Date, URL, Subject, Brief Summary, and What is Important. How to set up Credentials Setup:** Connect your Google Gemini and Google Sheets accounts via the n8n Credentials Manager. Google Sheet Linking:* In the *Append row in sheet and Append row in sheet in Google Sheets1 nodes, replace the **placeholder IDs and GIDs with the actual ID and sheet name of your dedicated UPSC notes spreadsheet. Scheduling:* Adjust the time in the *Schedule Trigger: Daily at 7 AM node** if you want the daily analysis to run at a different hour. AI Customization (Optional):* You can refine the System Message in the *AI Agent: Filter & Analyze UPSC News node** to focus the analysis on specific exam phases (e.g., Prelims only) or adjust the priority of subjects.
by Siddharth Gupta
Quick overview This workflow scans HTML files in a Google Drive folder, extracts and stores page text in Postgres, generates local vector embeddings with Ollama, and uses PGVector similarity searches to produce CSV reports that flag semantically duplicate website pages. How it works Starts manually and clears the existing PGVector embeddings table and the scraped page text table in Postgres. Lists files in a specified Google Drive folder, filters to the target documents, and processes them in batches. Downloads each HTML file from Google Drive, extracts the main body text, cleans it, and upserts the results into a Postgres table for scraped pages. Reads the scraped page text back from Postgres in batches, splits it into overlapping chunks, and attaches page metadata (sheet_id, file_name, file_url) to each chunk. Generates embeddings locally with Ollama and inserts the chunk vectors and metadata into Postgres (PGVector), deduplicating already-processed pages. Builds an HNSW index in Postgres, computes chunk-to-chunk similarity matches and a pairwise page report, and exports the results as a CSV file. Computes page-level centroid embeddings, finds highly similar page pairs, and exports a page-level duplicate report as a CSV file. Setup Add Google Drive OAuth2 credentials and set the Google Drive folder URL/ID used to scan for your HTML files. Add Postgres credentials for a database with the pgvector extension enabled and permissions to create/alter tables and indexes (including HNSW indexes). Add an Ollama credential and ensure the embedding model mxbai-embed-large:latest is available on your Ollama instance. Confirm your source files are HTML documents and that the workflow’s text extraction and similarity thresholds match your content and desired duplicate sensitivity. Requirements Working instance of n8n, either self-hosted or on the cloud. Remember, this workflow can be computationally expensive. Google Drive API (with OAuth setup in n8n credentials section) Ollama (for open source models) or any Embedding model API PostgreSQL with PGVector or any other vector database PgAdmin (for PostgreSQL) or your interface to access database tables via SQL for troubleshooting (optional). Additional info Limitations and Enhancements: Physical system memory mxbai-embed-large Running through Ollama is free and private, but the embedding generation speed depends entirely on your hardware. The more system memory you have, the more data you can process in batches in the loop node. Similarity threshold and boilerplate content The cosine distance used in this workflow is 0.15 for chunk-level matching. And 0.05 (similarity above 95%) of the threshold is used for page-level centroid matching. This is only the starting point. Once you have the data, and especially if your data has more noise, you might need to tweak these thresholds for better matching. This workflow needs HTML files to extract text This workflow doesn't crawl a website or fetch pages by entering a URL. You need to download HTML files (rendered or source) for consumption. Use parallel processing and Cloud APIs Two sub-processes take the most time: Downloading HTML files from Google Drive Creating vector embeddings If you can use parallel processing in n8n and execute these sub-processes in parallel, the process will be done much faster. Additionally, if you can use cloud APIs for embedding, it may save some you some processing time as well. Use efficient SQL queries Since I am from a non-tech background and not a coder, I used a mix of Gemini, Perplexity and Claude to create SQL codes for this workflow. If you're better at it, you can run computationally efficient queries that would help you achieve better results with less computation expense and time.
by vinci-king-01
Software Vulnerability Tracker with Pushover and Notion ⚠️ COMMUNITY TEMPLATE DISCLAIMER: This is a community-contributed template that uses ScrapeGraphAI (a community node). Please ensure you have the ScrapeGraphAI community node installed in your n8n instance before using this template. This workflow automatically scans multiple patent databases on a weekly schedule, filters new filings relevant to selected technology domains, saves the findings to Notion, and pushes instant alerts to your mobile device via Pushover. It is ideal for R&D teams and patent attorneys who need up-to-date insights on emerging technology trends and competitor activity. Pre-conditions/Requirements Prerequisites An n8n instance (self-hosted or n8n cloud) ScrapeGraphAI community node installed Active Notion account with an integration created Pushover account (user key & application token) List of technology keywords / CPC codes to monitor Required Credentials ScrapeGraphAI API Key** – Enables web scraping of patent portals Notion Credential** – Internal Integration Token with database write access Pushover Credential** – App Token + User Key for push notifications Additional Setup Requirements | Service | Needed Item | Where to obtain | |---------|-------------|-----------------| | USPTO, EPO, WIPO, etc. | Public URLs for search endpoints | Free/public | | Notion | Database with properties: Title, Abstract, URL, Date | Create in Notion | | Keyword List | Text file or environment variable PATENT_KEYWORDS | Define yourself | How it works This workflow automatically scans multiple patent databases on a weekly schedule, filters new filings relevant to selected technology domains, saves the findings to Notion, and pushes instant alerts to your mobile device via Pushover. It is ideal for R&D teams and patent attorneys who need up-to-date insights on emerging technology trends and competitor activity. Key Steps: Schedule Trigger**: Fires every week (default Monday 08:00 UTC). Code (Prepare Queries)**: Builds search URLs for each keyword and data source. SplitInBatches**: Processes one query at a time to respect rate limits. ScrapeGraphAI**: Scrapes patent titles, abstracts, links, and publication dates. Code (Normalize & Deduplicate)**: Cleans data, converts dates, and removes already-logged patents. IF Node**: Checks whether new patents were found. Notion Node**: Inserts new patent entries into the specified database. Pushover Node**: Sends a concise alert summarizing the new filings. Sticky Notes**: Document configuration tips inside the workflow. Set up steps Setup Time: 10-15 minutes Install ScrapeGraphAI: In n8n, go to “Settings → Community Nodes” and install @n8n-nodes/scrapegraphai. Add Credentials: ScrapeGraphAI: paste your API key. Notion: add the internal integration token and select your database. Pushover: provide your App Token and User Key. Configure Keywords: Open the first Code node and edit the keywords array (e.g., ["quantum computing", "Li-ion battery", "5G antenna"]). Point to Data Sources: In the same Code node, adjust the sources array if you want to add/remove patent portals. Set Notion Database Mapping: In the Notion node, map properties (Name, Abstract, Link, Date) to incoming JSON fields. Adjust Schedule (optional): Double-click the Schedule Trigger and change the CRON expression to your preferred interval. Test Run: Execute the workflow manually. Confirm that the Notion page is populated and a Pushover notification arrives. Activate: Switch the workflow to “Active” to enable automatic weekly execution. Node Descriptions Core Workflow Nodes: Schedule Trigger** – Defines the weekly execution time. Code (Build Search URLs)** – Dynamically constructs patent search URLs. SplitInBatches** – Sequentially feeds each query to the scraper. ScrapeGraphAI** – Extracts patent metadata from HTML pages. Code (Normalize Data)** – Formats dates, adds UUIDs, and checks for duplicates. IF** – Determines whether new patents exist before proceeding. Notion** – Writes new patent records to your Notion database. Pushover** – Sends real-time mobile/desktop notifications. Data Flow: Schedule Trigger → Code (Build Search URLs) → SplitInBatches → ScrapeGraphAI → Code (Normalize Data) → IF → Notion & Pushover Customization Examples Change Notification Message // Inside the Pushover node "Message" field return { message: 📜 ${items[0].json.count} new patent(s) detected in ${new Date().toDateString()}, title: '🆕 Patent Alert', url: items[0].json.firstPatentUrl, url_title: 'Open first patent' }; Add Slack Notification Instead of Pushover // Replace the Pushover node with a Slack node { text: ${$json.count} new patents published:\n${$json.list.join('\n')}, channel: '#patent-updates' } Data Output Format The workflow outputs structured JSON data: { "title": "Quantum Computing Device", "abstract": "A novel qubit architecture that ...", "url": "https://patents.example.com/US20240012345A1", "publicationDate": "2024-06-01", "source": "USPTO", "keywordsMatched": ["quantum computing"] } Troubleshooting Common Issues No data returned – Verify that search URLs are still valid and the ScrapeGraphAI selector matches the current page structure. Duplicate entries in Notion – Ensure the “Normalize Data” code correctly checks for existing URLs or IDs before insert. Performance Tips Limit the number of keywords or schedule the workflow during off-peak hours to reduce API throttling. Enable caching inside ScrapeGraphAI (if available) to minimize repeated requests. Pro Tips: Use environment variables (e.g., {{ $env.PATENT_KEYWORDS }}) to manage keyword lists without editing nodes. Chain an additional “HTTP Request → ML Model” step to auto-classify patents by CPC codes. Create a Notion view filtered by publicationDate is within past 30 days for quick scanning.
by Nguyen Thieu Toan
Monitor Facebook Pages and Analyze Content Safety via Telegram This n8n template automates the collection, storage, and safety analysis of Facebook posts while simultaneously providing an interactive AI assistant on Telegram. If you manage communities or brand pages and need to stay instantly informed about toxic content while having a smart assistant to answer quick operational queries, this workflow is perfect for you. How it works Interactive Chatbot (Trigger):* The Telegram Trigger listens for direct messages. An *AI Agent* (powered by Google Gemini) processes the input using *MongoDB* for conversation memory and custom tools (like *SerpAPI**) for deep research. Data Scraping (Schedule):* A Schedule Trigger runs every 3 hours to fetch the latest posts from your specified Facebook page using the *Apify Facebook Scraper**. Data Normalization & Storage:* Extracted posts are normalized and upserted into an *n8n Data Table**. This prevents duplicate processing of the same posts in future runs. Safety Analysis:** Post text and downloaded images are merged and sent to a secondary AI Agent. The AI evaluates the context and user reactions to flag the content as "Safe" or "Toxic". Smart Notification:** The safety report is beautifully formatted using Telegram HTML and dispatched directly to the admin's Telegram inbox. How to use Connect your Telegram Bot API credentials in both the Telegram Trigger and Send nodes. Connect your Google Gemini API key in all Language Model nodes. Connect your Apify API credentials and SerpAPI key. Configure your MongoDB connection for the chat memory nodes. Create an n8n Data Table (e.g., facebook_news_db) with a postId column (Number) and update the Data Table Upsert node to select your table. Customize the Set Context (Chat) and Set Context (Scraper) nodes with your specific details (Telegram Admin ID, Facebook Page URL, Bot Name). Activate the workflow and let the automation run. Requirements n8n Version:* Built and tested on *n8n 2.9.4+*. *(Note: You may encounter errors on older versions. It is highly recommended to update to the latest n8n version to use this workflow effectively). Google Gemini** API credentials. Telegram Bot** token. Apify** API credentials. SerpAPI** credentials. MongoDB** connection string. An active n8n Data Table. Customizing this workflow Change the scraper:** Swap the Apify node with any other social media scraping tool or RSS feed to monitor different platforms (e.g., X, LinkedIn). Change the database:** Replace the MongoDB Chat Memory node with Postgres or another memory node if you prefer a different database structure. Modify the AI persona:** Update the system prompt in the AI Agent nodes to change the chatbot's tone or the strictness of the safety evaluation. About the Author Created by: Nguyen Thieu Toan (Jay Nguyen) Email: me@nguyenthieutoan.com Website: nguyenthieutoan.com Company: GenStaff (genstaff.net) Socials (Facebook / X / LinkedIn): @nguyenthieutoan More templates: n8n.io/creators/nguyenthieutoan
by Atta
Never guess your SEO strategy again. This advanced workflow automates the most time-consuming part of SEO: auditing competitor articles and identifying exactly where your brand can outshine them. It extracts deep content from top-ranking URLs, compares it against your specific brand identity, and generates a ready-to-use "Action Plan" for your content team. The workflow uses Decodo for high-fidelity scraping, Gemini 2.5 Flash for strategic gap analysis, and Google Sheets as a dynamic "Brand Brain" and reporting dashboard. ✨ Key Features Brand-Centric Auditing:* Unlike generic SEO tools, this engine uses a live Google Sheet containing your *Brand Identity** to find "Content Gaps" specific to your unique value proposition. Automated SERP Itemization:** Converts a simple list of keywords into a filtered list of top-performing competitor URLs. Deep Markdown Extraction:** Uses Decodo Universal to bypass bot-blockers and extract clean Markdown content, preserving headers and structure for high-fidelity AI analysis. Structured Action Plans:** Outputs machine-readable JSON containing the competitor's H1, their "Winning Factor," and a 1-sentence "Checkmate" instruction for your writers. ⚙️ How it Works Data Foundation: The workflow triggers (Manual or Scheduled) and pulls your Global Config (e.g., result limits) and Brand Identity from a dedicated Google Sheet. Market Discovery: It retrieves your target keywords and uses the Decodo Google Search node to identify the top competitors. A Code Node then "itemizes" these results into individual URLs. Intelligence Harvesting: Decodo Universal scrapes each URL, and an HTML 5 node extracts the body content into Markdown format to minimize token noise for the AI. Strategic Audit: The AI Content Auditor (powered by Gemini) receives the competitor’s text and your Brand Identity. It identifies what the competitor missed that your brand excels at. Reporting Deck: The final Strategy Master Writer node appends the analysis—including the "Content Gap" and "Action Plan"—into a master Google Sheet for your marketing team. 📥 Component Installation This workflow relies on the Decodo node for search and scraping precision. Install Node: Click the + button in n8n, search for "Decodo," and add it to your canvas. Credentials: Use your Decodo API key. (Tip: Use a residential proxy setting for difficult sites like Reddit or Stripe). Gemini: Ensure you have the Google Gemini Chat Model node connected to the AI Agent. 🎁 Get a free Web Scraping API subscription here 👉🏻 https://visit.decodo.com/X4YBmy 🛠️ Setup Instructions 1. Google Sheets Configuration Create a spreadsheet with the following three tabs: Target Keywords**: One column named Target Keyword. Brand Identity**: One cell containing your brand mission, USPs, and target audience. Competitor Audit Feed**: Headers for Keyword, URL, Rank, Winning Factor, Content Gap, and Action Plan. Clone the spreadsheet here. 2. Global Configuration In the Config (Set) node, define your serp_results_amount (e.g., 10). This controls how many competitors are analyzed per keyword. ➕ How to Adapt the Template Competitor Exclusion:* Add a *Filter** node after "Market Discovery" to automatically skip domains like amazon.com or reddit.com if they aren't relevant to your niche. Slack Alerts:* Connect a *Slack** node after the AI analysis to notify your content manager immediately when a high-impact "Action Plan" is generated for a priority keyword. Multi-Model Verification:* Swap Gemini with *Claude 3.5 Sonnet* or *GPT-4o** in the Strategic Audit section to compare different AI perspectives on the same competitor content.
by Bhuvanesh R
Your Cold Email is Now Researched. This pipeline finds specific bottlenecks on prospect websites and instantly crafts an irresistible pitch 🎯 Problem Statement Traditional high-volume cold email outreach is stuck on generic personalization (e.g., "Love your website!"). Sales teams, especially those selling high-value AI Receptionists, struggle to efficiently find the one Unique Operational Hook (like manual scheduling dependency or high call volume) needed to make the pitch relevant. This forces reliance on expensive, slow manual research, leading to low reply rates and inefficient spending on bulk outreach tools. ✨ Solution This workflow deploys a resilient Dual-AI Personalization Pipeline that runs on a batch basis. It uses the Filter (Qualified Leads) node as a cost-saving Quality Gate to prevent processing bad leads. It executes a Targeted Deep Dive on successful leads, using GPT-4 for analytical insight extraction and Claude Sonnet for coherent, human-like copy generation. The entire process outputs campaign-ready data directly to Google Sheets and sends a critical QA Draft via Gmail. ⚙️ How It Works (Multi-Step Execution) 1\. Ingestion and Cost Control (The Quality Gate) Trigger and Ingestion:* The workflow starts via a *Manual Trigger, pulling leads directly from **Get All Leads (Google Sheets). Cost Filtering:* The *Filter (Qualified Leads)** node removes leads that lack a working email or website URL. Execution Isolation:* The *Loop Over Leads* node initiates individual processing. The *Capture Lead Data (Set)** node immediately captures and locks down the original lead context for stability throughout the loop. Hybrid Scraping:* The *Scrape Site (HTTP Request)* and *Extract Text & Links (HTML)* nodes execute the *Hybrid Scraping* strategy, simultaneously capturing *website text* and *external links**. Data Shaping & Status:* The *Filter Social & Status (Code)* node is the control center. It filters links, bundles the context, and critically, assigns a *status** of 'Success' or 'Scrape Fail'. Cost Control Branch:* The *If (IF node)* checks this status. Items with 'Scrape Fail' bypass all AI steps (saving *100% of AI token costs) and jump directly to **Log Final Result. Successful items proceed to the AI core. 2\. Dual-AI Coherence & Dispatch (The Executive Output) Analytical Synthesis:* The *Summarize Website (OpenAI)* node uses *GPT-4* to synthesize the full context and extract the *Unique Operational Hook** (e.g., manual booking overhead). Coherent Copy Generation:* The *Generate Subject & Body (Anthropic)* node uses the *Claude Sonnet* model to generate the subject and the multi-line body, guaranteeing *coherence** by creating both simultaneously in a single JSON output. Final Parsing:* The *Parse AI Output (Code)* node reliably strips markdown wrappers and extracts the clean *subject* and *body** strings. Final Delivery:* The data is logged via *Log Final Result (Google Sheets), and the completed email is sent to the user via **Create a draft (Gmail) for final Quality Assurance before sending. 🛠️ Setup Steps Before running the workflow, ensure these credentials and data structures are correctly configured: Credentials Anthropic:** Configure credentials for the Language Model (Claude Sonnet). OpenAI:** Configure credentials for the Analytical Model (GPT-4/GPT-4o). Google Services:* Set up OAuth2 credentials for *Google Sheets* (Input/Output) and *Gmail** (Draft QA and Completion Alert). Configuration Google Sheet Setup:* Your input sheet must include the columns *email, **website\_url, and an empty Icebreaker column for initial filtering. HTTP URL:* Verify that the *Scrape Site** node's URL parameter is set to pull the website URL from the stabilized data structure: ={{ $json.website\_url }}. AI Prompts:** Ensure the Anthropic prompt contains your current Irresistible Sales Offer and the required nested JSON output structure. ✅ Benefits Coherence Guarantee:* A single *Anthropic** node generates both the subject and body, guaranteeing the message is perfectly aligned and hits the same unique insight. Maximum Cost Control:* The *IF node* prevents spending tokens on bad or broken websites, making the campaign highly *budget-efficient**. Deep Personalization:* Combines *website text* and *social media links**, creating an icebreaker that implies thorough, manual research. High Reliability:* Uses robust *Code nodes** for data structuring and parsing, ensuring the workflow runs consistently under real-world conditions without crashing. Zero-Risk QA:* The final *Gmail (Create a draft)** step ensures human review of the generated copy before any cold emails are sent out.
by phil
This workflow is designed for B2B professionals to automatically identify and summarize business opportunities from a company's website. By leveraging Bright Data's Web Unblocker and advanced AI models from OpenRouter, it scrapes relevant company pages ("About Us", "Team", "Contact"), analyzes the content for potential pain points and needs, and synthesizes a concise, actionable report. The final output is formatted for direct use in documents, making it an ideal tool for sales, marketing, and business development teams to prepare for prospecting calls or personalize outreach. Who's it for This template is ideal for: B2B Sales Teams:** Quickly find and qualify leads by identifying specific business needs before a cold call. Marketing Agencies:** Develop personalized content and value propositions based on a prospect's public website information. Business Development Professionals:** Efficiently research potential partners or clients and discover collaboration opportunities. Entrepreneurs:** Gain a competitive edge by understanding a competitor's strategy or a potential client's operations. How it works The workflow is triggered by a chat message, typically a URL from an n8n chat application. It uses Bright Data to scrape the website's sitemap and extract all anchor links from the homepage. An AI agent analyzes the extracted URLs to filter for pages relevant to company information (e.g., "about-us," "team," "contact"). The workflow then scrapes the content of these specific pages. A second AI agent summarizes the content of each page, looking for business opportunities related to AI-powered automation. The summaries are merged and a final AI agent synthesizes them into a single, cohesive report, formatted for easy reading in a Google Doc. How to set up Bright Data Credentials: Sign up for a Bright Data account and create a Web Unblocker zone. In n8n, create new Bright Data API credentials and copy your API key. OpenRouter Credentials: Create an account on OpenRouter and get your API key. In n8n, create new OpenRouter API credentials and paste your key. Chat Trigger Node: Configure the "When chat message received" node. Copy the production webhook URL to integrate with your preferred chat platform. Requirements An active n8n instance. A Bright Data account with a Web Unblocker zone. An OpenRouter account with API access. How to customize this workflow AI Prompting:** Edit the "systemMessage" parameters in the "AI Agent", "AI Agent1", and "AI Agent2" nodes to change the focus of the opportunity analysis. For example, modify the prompts to search for specific technologies, industry jargon, or different types of business challenges. Model Selection:** The workflow uses openai/o4-mini and openai/gpt-5. You can change these to other models available on OpenRouter by editing the model parameter in the OpenRouter Chat Model nodes. Scraping Logic:** The extract url node uses a regular expression to find `` tags. This can be modified or replaced with an HTML Extraction node to target different elements or content on a website. Output Format:** The final output is designed for Google Docs. You can modify the last "AI Agent2" node's prompt to generate the output in a different format, such as a simple JSON object or a markdown list. Phil | Inforeole 🇫🇷 Contactez nous pour automatiser vos processus
by Davide
This workflow automates the entire process of collecting, analyzing, and reporting customer reviews from Feedaty (similar to Trustpilot) using ScrapeGraphAI, transforming raw user feedback into a structured, management-ready reputation report in PDF using new Gemini 3 model and ConvertAPI & Upload to Google Drive. Key Advantages ✅ End-to-End Automation From data collection to final PDF delivery, the entire reputation analysis process is fully automated, eliminating manual scraping, copy-paste work, and reporting overhead. ✅ AI-Driven, Management-Ready Insights The workflow does not just summarize reviews it interprets them strategically, producing insights that are immediately useful for: Management Marketing Customer Support Operations Product & UX teams ✅ Structured & Consistent Reporting Every execution produces reports with the same structure, metrics, and logic, making it ideal for: Periodic reputation monitoring Trend analysis over time Internal performance reviews ✅ Scalable & Configurable Easily adaptable to any Feedaty company profile Page limits and review volume can be adjusted without changing logic Can be scheduled or extended to multiple brands ✅ Data Quality & Compliance No personal data exposure Explicit handling of missing or ambiguous information No assumptions or hallucinated insights Fully transparent and audit-friendly output ✅ Seamless Stakeholder Distribution Automatic upload to Google Drive ensures reports are centralized, shareable, and accessible, with no additional manual steps. Ideal Use Cases Brand & reputation monitoring Customer experience audits Quarterly or monthly executive reports Pre-sales or investor documentation Customer support performance evaluation How it works This workflow automates the entire process of collecting, analyzing, and reporting customer feedback from Feedaty. It starts by scraping live reviews from a specified company's Feedaty page using ScrapeGraphAI, extracting review details like date, rating, and text. Each review is then individually analyzed for sentiment (Positive, Neutral, or Negative) using an AI model. All processed reviews are aggregated and passed to a specialized AI agent that performs a comprehensive company-level reputation analysis, generating a structured management report. Finally, the report is converted into an HTML/PDF format and uploaded to a designated Google Drive folder, creating a fully automated pipeline from data collection to actionable insights delivery. Set up steps Configure Parameters: Set the Feedaty company identifier (e.g., maxisport) and the maximum number of review pages to scrape in the "Set Parameters" node. API Credentials: Ensure the following credentials are configured in n8n: ScrapeGraphAI API (for web scraping) Google Gemini API (for AI sentiment analysis and report generation) Google Drive OAuth2 (for file upload) ConvertAPI (for HTML to PDF conversion) Customize Output: Optionally adjust the "Limit reviews" node to control the number of reviews processed and modify the AI agent's system prompt in "Company Reputation Management" to tailor the report format. Destination Folder: Verify the Google Drive folder ID in the "Upload file" node points to the correct destination for the generated reports. Execution: Trigger the workflow manually via the "When clicking ‘Test workflow’" node to run the complete scraping, analysis, and reporting pipeline. 👉 Subscribe to my new YouTube channel. Here I’ll share videos and Shorts with practical tutorials and FREE templates for n8n. Need help customizing? Contact me for consulting and support or add me on Linkedin.
by iamvaar
Quick overview Youtube Video: https://youtu.be/h47YeiivZvs?si=V-S1X-ME5V24Tuch This workflow runs daily to pull the latest Product Hunt posts, qualifies them with Google Gemini, scrapes ICP-matching product websites with Apify to find emails, then writes the lead data to Google Sheets and creates contacts in GoHighLevel. How it works Runs every day on a schedule trigger. Queries the Product Hunt GraphQL API for the latest posts and iterates through each result. Resolves each product’s website redirect target, extracts a clean domain, and filters out blocked domains like app stores, social networks, and link aggregators. Sends the product name, tagline, and description to a Google Gemini–powered agent to score ICP fit and keeps only matches. Uses Apify to crawl each matched website and extract email addresses from the site content. Groups scraped results by domain, keeps only products with at least one email, and appends or updates a matching row in Google Sheets. Creates a GoHighLevel contact using the first email and stores any additional emails in a custom field. Setup Add Product Hunt OAuth2 credentials and ensure the GraphQL query in the Product Hunt request returns the fields used (name, tagline, description, website, topics, makers). Add Google Gemini (Google PaLM) credentials for the AI qualification step. Add an Apify API token and set the actor ID and crawl settings to match your email-scraping needs. Connect Google Sheets with a service account, set the target spreadsheet and sheet, and ensure the sheet has columns for product_name, tagline, description, makers, url, original_ph_url, and emails. Add GoHighLevel OAuth2 credentials and update the contact field mappings, including the custom field ID used to store additional emails.