Skip to content

CrawlForge TeamEngineering Team

8 min read

Scrape Amazon, LinkedIn & 8 More Sites With One Tool

Update (30 August 2026): the linkedin-profile and tweet templates described below have been retired. LinkedIn's robots.txt disallows every path for all agents but its own crawler and profiles sit behind an authentication wall; X's robots.txt disallows every path for generic agents and its keyless embed endpoints are disallowed by their own robots.txt. Naming either template now returns the reason and fetches nothing. reddit-thread reads the post from the Arctic Shift archive instead of reddit.com, and reddit_search reads the comment tree.

Half the scraping requests we see at CrawlForge are the same ten sites: Amazon, LinkedIn, GitHub, YouTube, Reddit, Hacker News, Stack Overflow, npm, Product Hunt, and Twitter/X. We got tired of watching people write the same CSS selectors over and over -- and watching those selectors break the next time the site updated its layout. So we did the work once, packaged it as scrape_template, and now you pay 1 credit and get structured JSON.

Table of Contents

What Is scrape_template?

scrape_template is a single CrawlForge tool with ten pre-built site schemas. You pick the template, pass a URL, and get back structured JSON matching that site's natural shape. No CSS selectors. No HTML parsing. No schema definition.

The trade-off: you only get the ten sites we maintain. If you need something else, use scrape_structured (CSS-first) or extract_with_llm (LLM-first). For the long tail of "I want product data from Amazon" requests, scrape_template is the shortest path. Need a multi-step workflow instead of a single site? See how to use the templates gallery.

It costs 1 credit per scrape -- the same as a basic fetch_url -- because we have already done the schema work upstream.

The 10 Supported Sites

TemplateReturnsBest forExample URL pattern
amazon-productTitle, price, rating, review count, images, ASIN, availabilityPrice monitoring, product research/dp/<ASIN>
linkedin-profileName, headline, location, about, current companyLead enrichment/in/<handle>
github-repoStars, forks, language, topics, license, last updatedRepo analysis, AI training data/<owner>/<repo>
youtube-videoTitle, channel, views, duration, published, descriptionContent research/watch?v=<id>
reddit-threadPost title, score, author, subreddit, bodyCommunity signals/r/<sub>/comments/<id>
hacker-news-front-pageFront-page stories: title, URL, score, author, commentsTech trend trackingnews.ycombinator.com
stackoverflow-questionQuestion, accepted answer, vote counts, tagsDeveloper Q&A mining/questions/<id>
npm-packagePackage metadata, weekly downloads, version, maintainersDependency analysis/package/<name>
producthunt-launchProduct, tagline, upvotes, topics, websiteLaunch monitoring/posts/<slug>
tweetText, author, URL, imageSocial listening/<user>/status/<id>

Quick Start: Scrape an Amazon Product

Bash
crawlforge template amazon-product "https://www.amazon.com/dp/B0CHX1W1XY"

Output:

Json
{
  "asin": "B0CHX1W1XY",
  "title": "Logitech MX Master 3S Wireless Performance Mouse",
  "price": { "amount": 99.99, "currency": "USD" },
  "rating": 4.7,
  "review_count": 12483,
  "in_stock": true,
  "images": ["https://m.media-amazon.com/...", "..."],
  "credits_used": 1
}

From an MCP client like Claude Code:

"Use scrape_template with the amazon template to get the current price and rating for ASIN B0CHX1W1XY."

Claude picks the tool, formats the call, and returns the data. One credit.

Bash
crawlforge template linkedin-profile "https://www.linkedin.com/in/satyanadella"

Output:

Json
{
  "name": "Satya Nadella",
  "headline": "Chairman and CEO at Microsoft",
  "location": "Redmond, Washington",
  "current_role": { "title": "CEO", "company": "Microsoft", "since": "2014-02" },
  "experience_count": 6,
  "skills_top": ["Leadership", "Strategy", "Cloud Computing"],
  "credits_used": 1
}

A note on LinkedIn scraping. LinkedIn's terms of service restrict automated access. The hiQ Labs v. LinkedIn case (9th Circuit, 2022) established that scraping public profile data is generally permissible, but commercial use, login-required scraping, and aggressive frequency can still trigger legal action and ToS bans. Use scrape_template with the linkedin-profile template for public, low-frequency, non-resold data only.

GitHub Repos for AI Training Data

Bash
crawlforge template github-repo "https://github.com/anthropics/anthropic-sdk-python"

Output:

Json
{
  "owner": "anthropics",
  "name": "anthropic-sdk-python",
  "stars": 1842,
  "forks": 287,
  "primary_language": "Python",
  "languages": { "Python": 98.4, "Makefile": 1.6 },
  "license": "MIT",
  "topics": ["claude", "anthropic", "sdk"],
  "readme_markdown": "# Anthropic Python SDK...",
  "last_commit_at": "2026-05-19T14:22:11Z",
  "credits_used": 1
}

This template is heavily used for AI training-data pipelines -- pulling READMEs at scale across thousands of repos. Pair it with batch_scrape to process a CSV of repo URLs.

The Other Seven Templates

YouTube -- title, channel, views, full transcript when available:

Bash
crawlforge template youtube-video "https://www.youtube.com/watch?v=dQw4w9WgXcQ"

Reddit -- post + comment tree:

Bash
crawlforge template reddit-thread "https://www.reddit.com/r/programming/comments/<id>"

Hacker News -- the front page as a list of stories:

Bash
crawlforge template hacker-news-front-page "https://news.ycombinator.com"
# returns up to 30 front-page stories; slice the top 10 with jq:
crawlforge template hacker-news-front-page "https://news.ycombinator.com" --json | jq '.stories[:10]'

Stack Overflow -- question, accepted answer, top alternatives:

Bash
crawlforge template stackoverflow-question "https://stackoverflow.com/questions/12345678"

npm -- package metadata + weekly downloads:

Bash
crawlforge template npm-package "https://www.npmjs.com/package/next"

Product Hunt -- product, makers, upvotes:

Bash
crawlforge template producthunt-launch "https://www.producthunt.com/posts/crawlforge"

Twitter/X -- single tweet with engagement and replies:

Bash
crawlforge template tweet "https://x.com/elonmusk/status/<id>"

All return JSON. All cost 1 credit. All maintained centrally -- when LinkedIn or Amazon updates their layout, we update the template.

scrape_template vs scrape_structured vs extract_with_llm

A decision tree:

Is your target one of the 10 supported sites? Yes -> use scrape_template (1 credit, maintained for you) No Do you know the CSS selectors and are they stable? Yes -> use scrape_structured (2 credits, you maintain selectors) No -> use extract_with_llm (3 credits, schema-based, layout-resilient)

Quick comparison:

scrape_templatescrape_structuredextract_with_llm
Credits123
Coverage10 specific sitesAny site you can write selectors forAny site
MaintenanceWe maintainYou maintainLLM adapts
SpeedFast (cached schemas)FastSlower (LLM call)
Best forPopular sites, high volumeSpecific known structureUnknown or shifting structure

Limitations

  • Only 10 sites. If you need Etsy, eBay, TikTok, or others, you are waiting on the roadmap or rolling your own with scrape_structured / extract_with_llm. Request templates on Discord.
  • Public data only. No template requires login. Profiles set to private, gated repos, and protected tweets will return what is publicly visible only.
  • Layout changes happen. When a site ships a redesign, we usually have the template patched within 24 hours.
  • Rate limits apply. Heavy-volume LinkedIn or Amazon scraping should pair scrape_template with stealth_mode (5 credits) and respect each site's robots.txt.

Ready to skip the selectors? Start free with 1,000 credits -- enough for 1,000 template scrapes. New here? Read the v4.2.2 launch post for context, or the e-commerce extraction guide for a real-world workflow built around these templates.

Try this yourself — no signup needed

Explore all 31 CrawlForge scraping and extraction tools in the playground, then start free with 1,000 credits.

1,000 free credits • One-time • No credit card required

Tags

  • scrape-template
  • Amazon
  • LinkedIn
  • GitHub
  • use-cases
  • pre-built-scrapers

About the Author

CrawlForge Team

Engineering Team

Building the most comprehensive web scraping MCP server. We create tools that help developers extract, analyze, and transform web data for AI applications.

Newsletter

Stay updated with the latest insights

Get tutorials, product updates, and web scraping tips delivered to your inbox.

No spam. Unsubscribe anytime.

FAQ

Frequently asked questions

01What sites does scrape_template support?

Ten sites in v4.2.2: Amazon, LinkedIn, GitHub, YouTube, Reddit, Hacker News, Stack Overflow, npm, Product Hunt, and Twitter/X. Each has a pre-built schema returning the fields you would normally want (product price/rating, profile name/role, repo stars/README, video transcript, etc.). More templates are coming in v4.3.

02Is scraping LinkedIn legal?

The hiQ Labs v. LinkedIn case (9th Circuit, 2022) established that scraping public profile data is generally permissible, but LinkedIn's ToS restricts automated access -- and aggressive scraping or commercial resale can still trigger legal action and bans. Use scrape_template with the linkedin-profile template for public, low-frequency, non-resold use cases. Consult a lawyer if you are scraping at scale or for commercial products.

03Can I add a custom template?

Not directly today, but we accept template requests on Discord and prioritize by demand. Sites with significant request volume (Etsy, eBay, TikTok, Instagram, Google Maps) are on the roadmap for v4.3. For one-off custom work, use scrape_structured (CSS selectors) or extract_with_llm (schema-driven).

04What is the difference between scrape_template and scrape_structured?

scrape_template is for ten specific sites where we already maintain the schema -- you just pick the template name. scrape_structured is general-purpose: you provide CSS selectors for any site, and CrawlForge runs them. Template is faster and cheaper (1 credit vs 2) when your target is one of the ten supported sites.

05How fresh are the scrape_template schemas?

We monitor each supported site for layout changes and typically ship a template patch within 24 hours of any breaking change. Updates are transparent to your code -- you keep calling the same template name and the data shape stays the same. If you notice a regression, report it on Discord or GitHub.

06What happens if a supported site changes its layout?

Calls keep returning JSON in the documented shape, even if the underlying selectors needed to change. We absorb the maintenance burden so you do not have to. If a layout change is severe enough to temporarily break a field, we mark that field nullable in the response until the patch is live (usually within 24 hours).

Keep reading

Related Articles

Use Cases

12m

Web Scraping by Industry: 2026 Playbook

Industry-specific web scraping strategies for real estate, finance, e-commerce, healthcare, and travel. Data targets, CrawlForge tools, and compliance rules.