Maton

Firecrawl

Access the Firecrawl API with managed authentication. Scrape webpages, crawl entire websites, map site URLs, and search the web with full content extraction.

Reference

Scrape

POST /firecrawl/v2/scrape

Extract content from a single webpage.

Required Parameters:

  • url (string): The webpage URL to scrape

Optional Parameters:

  • formats (array): Output formats - "markdown", "html", "json", "screenshot", "links" (default: ["markdown"])
  • onlyMainContent (boolean): Extract only main content, exclude headers/footers (default: true)
  • includeTags (array): HTML tags to include
  • excludeTags (array): HTML tags to exclude
  • waitFor (integer): Milliseconds to wait before scraping (default: 0)
  • timeout (integer): Request timeout in ms (default: 30000, max: 300000)
  • mobile (boolean): Emulate mobile device (default: false)
  • actions (array): Browser actions to perform before scraping
  • headers (object): Custom HTTP headers
  • blockAds (boolean): Block ads and cookie banners (default: true)

Example:

python <<'EOF'
import urllib.request, os, json
data = json.dumps({
    "url": "https://docs.firecrawl.dev",
    "formats": ["markdown", "html"],
    "onlyMainContent": True,
    "waitFor": 1000
}).encode()
req = urllib.request.Request('https://api.maton.ai/firecrawl/v2/scrape', data=data, method='POST')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
req.add_header('Content-Type', 'application/json')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Response:

{
  "success": true,
  "data": {
    "markdown": "# Example Domain\n\nThis domain is for use in documentation...",
    "metadata": {
      "title": "Example Domain",
      "language": "en",
      "sourceURL": "https://example.com",
      "url": "https://example.com/",
      "statusCode": 200,
      "contentType": "text/html",
      "creditsUsed": 1
    }
  }
}

Crawl (Start)

POST /firecrawl/v2/crawl

Start crawling an entire website. Returns a crawl ID for status polling.

Required Parameters:

  • url (string): The base URL to start crawling from

Optional Parameters:

  • limit (integer): Maximum pages to crawl (default: 10000)
  • maxDepth (integer): Maximum crawl depth
  • includePaths (array): Regex patterns for URLs to include
  • excludePaths (array): Regex patterns for URLs to exclude
  • allowSubdomains (boolean): Enable subdomain crawling
  • allowExternalLinks (boolean): Follow external links
  • scrapeOptions (object): Options for each page scrape (formats, onlyMainContent, etc.)
  • webhook (string): Webhook URL for completion notification

Example:

python <<'EOF'
import urllib.request, os, json
data = json.dumps({
    "url": "https://example.com",
    "limit": 10,
    "scrapeOptions": {
        "formats": ["markdown"]
    }
}).encode()
req = urllib.request.Request('https://api.maton.ai/firecrawl/v2/crawl', data=data, method='POST')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
req.add_header('Content-Type', 'application/json')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Response:

{
  "success": true,
  "id": "019cdc53-0acf-76ec-a80c-3ead753b2730",
  "url": "https://api.firecrawl.dev/v1/crawl/019cdc53-0acf-76ec-a80c-3ead753b2730"
}

Crawl (Get Status)

GET /firecrawl/v2/crawl/{id}

Get the status and results of a crawl job.

Path Parameters:

  • id (string): The crawl job ID

Example:

python <<'EOF'
import urllib.request, os, json
crawl_id = "019cdc53-0acf-76ec-a80c-3ead753b2730"
req = urllib.request.Request(f'https://api.maton.ai/firecrawl/v2/crawl/{crawl_id}')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Response:

{
  "success": true,
  "status": "completed",
  "completed": 2,
  "total": 2,
  "creditsUsed": 2,
  "expiresAt": "2026-03-12T09:56:00.000Z",
  "data": [
    {
      "markdown": "# Example Domain\n\nThis domain is for use in documentation...",
      "metadata": {
        "title": "Example Domain",
        "sourceURL": "https://example.com",
        "statusCode": 200
      }
    }
  ]
}

Status Values:

  • scraping - Crawl in progress
  • completed - Crawl finished successfully
  • failed - Crawl failed

Crawl (Cancel)

DELETE /firecrawl/v2/crawl/{id}

Cancel an in-progress crawl job.

Path Parameters:

  • id (string): The crawl job ID

Example:

python <<'EOF'
import urllib.request, os, json
crawl_id = "019cdc53-0acf-76ec-a80c-3ead753b2730"
req = urllib.request.Request(f'https://api.maton.ai/firecrawl/v2/crawl/{crawl_id}', method='DELETE')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Response:

{
  "success": true,
  "status": "cancelled"
}

Map

POST /firecrawl/v2/map

Get all URLs from a website without scraping content.

Required Parameters:

  • url (string): The starting URL

Optional Parameters:

  • search (string): Query to order results by relevance
  • limit (integer): Maximum links to return (default: 5000, max: 100000)
  • includeSubdomains (boolean): Include subdomains (default: true)
  • sitemap (string): Sitemap handling - "skip", "include", "only" (default: "include")
  • ignoreQueryParameters (boolean): Exclude URLs with query params (default: true)
  • timeout (integer): Timeout in milliseconds

Example:

python <<'EOF'
import urllib.request, os, json
data = json.dumps({
    "url": "https://docs.firecrawl.dev",
    "limit": 100,
    "includeSubdomains": False
}).encode()
req = urllib.request.Request('https://api.maton.ai/firecrawl/v2/map', data=data, method='POST')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
req.add_header('Content-Type', 'application/json')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Response:

{
  "success": true,
  "links": [
    "https://docs.firecrawl.dev",
    "https://docs.firecrawl.dev/api-reference",
    "https://docs.firecrawl.dev/introduction"
  ]
}
POST /firecrawl/v2/search

Search the web and get full page content for each result.

Required Parameters:

  • query (string): Search query (max 500 characters)

Optional Parameters:

  • limit (integer): Number of results (default: 5, max: 100)
  • sources (array): Search types - "web", "images", "news" (default: ["web"])
  • country (string): ISO country code (default: "US")
  • location (string): Geographic targeting (e.g., "Germany")
  • tbs (string): Time filter - "qdr:d" (day), "qdr:w" (week), "qdr:m" (month), "qdr:y" (year)
  • timeout (integer): Timeout in ms (default: 60000)
  • scrapeOptions (object): Options for content extraction

Example:

python <<'EOF'
import urllib.request, os, json
data = json.dumps({
    "query": "web scraping best practices",
    "limit": 5,
    "scrapeOptions": {
        "formats": ["markdown"]
    }
}).encode()
req = urllib.request.Request('https://api.maton.ai/firecrawl/v2/search', data=data, method='POST')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
req.add_header('Content-Type', 'application/json')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Response:

{
  "success": true,
  "data": [
    {
      "url": "https://example.com/article",
      "title": "Web Scraping Best Practices",
      "description": "Learn the best practices for web scraping...",
      "markdown": "# Web Scraping Best Practices\n\n..."
    }
  ],
  "creditsUsed": 5
}

Batch Scrape (Start)

POST /firecrawl/v2/batch/scrape

Scrape multiple URLs in a single batch job.

Required Parameters:

  • urls (array): List of URLs to scrape

Optional Parameters:

  • formats (array): Output formats (default: ["markdown"])
  • onlyMainContent (boolean): Extract only main content (default: true)
  • webhook (string): Webhook URL for completion notification

Example:

python <<'EOF'
import urllib.request, os, json
data = json.dumps({
    "urls": ["https://example.com", "https://example.org"],
    "formats": ["markdown"]
}).encode()
req = urllib.request.Request('https://api.maton.ai/firecrawl/v2/batch/scrape', data=data, method='POST')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
req.add_header('Content-Type', 'application/json')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Response:

{
  "success": true,
  "id": "019cdc59-56b9-7096-a9f9-95fcc92a3a75",
  "url": "https://api.firecrawl.dev/v1/batch/scrape/019cdc59-56b9-7096-a9f9-95fcc92a3a75"
}

Batch Scrape (Get Status)

GET /firecrawl/v2/batch/scrape/{id}

Get the status and results of a batch scrape job.

Path Parameters:

  • id (string): The batch scrape job ID

Example:

python <<'EOF'
import urllib.request, os, json
batch_id = "019cdc59-56b9-7096-a9f9-95fcc92a3a75"
req = urllib.request.Request(f'https://api.maton.ai/firecrawl/v2/batch/scrape/{batch_id}')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Response:

{
  "success": true,
  "status": "completed",
  "completed": 2,
  "total": 2,
  "creditsUsed": 2,
  "expiresAt": "2026-03-12T10:02:54.000Z",
  "data": [
    {
      "markdown": "# Example Domain\n\n...",
      "metadata": {
        "title": "Example Domain",
        "sourceURL": "https://example.com",
        "statusCode": 200
      }
    }
  ]
}

Batch Scrape (Cancel)

DELETE /firecrawl/v2/batch/scrape/{id}

Cancel an in-progress batch scrape job.

Path Parameters:

  • id (string): The batch scrape job ID

Batch Scrape (Get Errors)

GET /firecrawl/v2/batch/scrape/{id}/errors

Get errors from a batch scrape job.

Path Parameters:

  • id (string): The batch scrape job ID

Response:

{
  "errors": [],
  "robotsBlocked": []
}

Crawl (Get Errors)

GET /firecrawl/v2/crawl/{id}/errors

Get errors from a crawl job.

Path Parameters:

  • id (string): The crawl job ID

Example:

python <<'EOF'
import urllib.request, os, json
crawl_id = "019cdc53-0acf-76ec-a80c-3ead753b2730"
req = urllib.request.Request(f'https://api.maton.ai/firecrawl/v2/crawl/{crawl_id}/errors')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Response:

{
  "errors": [],
  "robotsBlocked": []
}

Crawl (Get Active)

GET /firecrawl/v2/crawl/active

Get all active crawl jobs.

Example:

python <<'EOF'
import urllib.request, os, json
req = urllib.request.Request('https://api.maton.ai/firecrawl/v2/crawl/active')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Response:

{
  "success": true,
  "crawls": []
}

Extract (Start)

POST /firecrawl/v2/extract

Extract structured data from URLs using AI.

Required Parameters:

  • urls (array): List of URLs to extract from
  • prompt (string): Natural language description of what to extract

Optional Parameters:

  • schema (object): JSON schema for structured output
  • scrapeOptions (object): Options for scraping

Example:

python <<'EOF'
import urllib.request, os, json
data = json.dumps({
    "urls": ["https://example.com"],
    "prompt": "Extract the main heading and description"
}).encode()
req = urllib.request.Request('https://api.maton.ai/firecrawl/v2/extract', data=data, method='POST')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
req.add_header('Content-Type', 'application/json')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Response:

{
  "success": true,
  "id": "019cdc59-977b-774b-b584-af2af45c055b",
  "urlTrace": []
}

Extract (Get Status)

GET /firecrawl/v2/extract/{id}

Get the status and results of an extract job.

Path Parameters:

  • id (string): The extract job ID

Example:

python <<'EOF'
import urllib.request, os, json
extract_id = "019cdc59-977b-774b-b584-af2af45c055b"
req = urllib.request.Request(f'https://api.maton.ai/firecrawl/v2/extract/{extract_id}')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Response:

{
  "success": true,
  "data": [
    {
      "heading": "Example Domain",
      "description": "This domain is for use in documentation..."
    }
  ],
  "status": "completed",
  "expiresAt": "2026-03-11T16:03:05.000Z"
}

Browser (Create Session)

POST /firecrawl/v2/browser

Create an interactive browser session for manual control via CDP.

Example:

python <<'EOF'
import urllib.request, os, json
data = json.dumps({}).encode()
req = urllib.request.Request('https://api.maton.ai/firecrawl/v2/browser', data=data, method='POST')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
req.add_header('Content-Type', 'application/json')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Response:

{
  "success": true,
  "id": "019cdc5d-5c9d-732e-a7bd-f095a96a2bb1",
  "cdpUrl": "wss://browser.firecrawl.dev/cdp/...",
  "liveViewUrl": "https://liveview.firecrawl.dev/...",
  "interactiveLiveViewUrl": "https://liveview.firecrawl.dev/...",
  "expiresAt": "2026-03-11T10:17:12.409Z"
}

Browser (List Sessions)

GET /firecrawl/v2/browser

List all active browser sessions.

Example:

python <<'EOF'
import urllib.request, os, json
req = urllib.request.Request('https://api.maton.ai/firecrawl/v2/browser')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Response:

{
  "success": true,
  "sessions": [
    {
      "id": "019cdc5d-5c9d-732e-a7bd-f095a96a2bb1",
      "status": "active",
      "cdpUrl": "wss://browser.firecrawl.dev/cdp/...",
      "liveViewUrl": "https://liveview.firecrawl.dev/..."
    }
  ]
}

Browser (Delete Session)

DELETE /firecrawl/v2/browser/{id}

Delete a browser session.

Path Parameters:

  • id (string): The browser session ID

Agent (Start)

POST /firecrawl/v2/agent

Start an AI agent to autonomously navigate and extract data.

Required Parameters:

  • prompt (string): Description of what data to extract (max 10,000 chars)

Optional Parameters:

  • urls (array): URLs to constrain the agent to
  • schema (object): JSON schema for structured output
  • maxCredits (integer): Maximum credits to use (default: 2500)
  • strictConstrainToURLs (boolean): Only visit provided URLs
  • model (string): "spark-1-mini" (default, cheaper) or "spark-1-pro" (higher accuracy)

Example:

python <<'EOF'
import urllib.request, os, json
data = json.dumps({
    "prompt": "Find the pricing information",
    "urls": ["https://example.com"],
    "model": "spark-1-mini"
}).encode()
req = urllib.request.Request('https://api.maton.ai/firecrawl/v2/agent', data=data, method='POST')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
req.add_header('Content-Type', 'application/json')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Response:

{
  "success": true,
  "id": "019cdc5d-a2d4-728c-9c91-e9eae475568f"
}

Agent (Get Status)

GET /firecrawl/v2/agent/{id}

Get the status and results of an agent job.

Path Parameters:

  • id (string): The agent job ID

Example:

python <<'EOF'
import urllib.request, os, json
agent_id = "019cdc5d-a2d4-728c-9c91-e9eae475568f"
req = urllib.request.Request(f'https://api.maton.ai/firecrawl/v2/agent/{agent_id}')
req.add_header('Authorization', f'Bearer {os.environ["MATON_API_KEY"]}')
print(json.dumps(json.load(urllib.request.urlopen(req)), indent=2))
EOF

Response:

{
  "success": true,
  "status": "completed",
  "model": "spark-1-pro",
  "data": {...},
  "expiresAt": "2026-03-12T10:07:30.055Z"
}

Agent (Cancel)

DELETE /firecrawl/v2/agent/{id}

Cancel an in-progress agent job.

Path Parameters:

  • id (string): The agent job ID

Resources

On this page