Skip to content

Result Polling

When scraping tasks take longer than the timeout limit or when you submit async requests, you’ll need to poll for results. This guide covers how to retrieve results from background tasks.

When Polling Is Needed Permalink to When Polling Is Needed

Polling is required in these scenarios:

  1. Explicit async requests β€” You set async=true in your request
  2. Request timeouts β€” A synchronous request exceeds the timeout limit and auto-converts to async

Timeout Handling Permalink to Timeout Handling

Different modes have different timeout limits:

Mode Timeout What Happens After
request 30 seconds Auto-converts to async task
browser 45 seconds Auto-converts to async task
auto 30-45 seconds Depends on mode used

Timeout Response Permalink to Timeout Response

When a synchronous request times out, you receive a 202 Accepted response:

{
  "success": true,
  "task_id": "task_abc123",
  "status": "processing",
  "message": "Task is taking longer than expected. Use task_id to check status.",
  "check_url": "/api/v1/scraper/tasks/task_abc123",
  "created_at": "2026-09-28T20:31:12.482Z",
  "started_at": "2026-09-28T20:31:40.101Z",
  "finished_at": null,
  "elapsed_seconds": 45.3
}

Submitting Async Requests Permalink to Submitting Async Requests

For long-running or batch jobs, submit async requests from the start:

Request Permalink to Request

curl -X POST "https://scrape.evomi.com/api/v1/scraper/realtime" \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com",
    "async": true
  }'

Response Permalink to Response

{
  "task_id": "task_abc123",
  "status": "processing",
  "check_url": "/api/v1/scraper/tasks/task_abc123",
  "created_at": "2026-09-28T20:31:12.482Z",
  "started_at": null,
  "finished_at": null,
  "elapsed_seconds": 0.1
}

Checking Task Status Permalink to Checking Task Status

Request Permalink to Request

curl "https://scrape.evomi.com/api/v1/scraper/tasks/task_abc123?api_key=YOUR_API_KEY"

Response (Processing β€” HTTP 202) Permalink to Response (Processing β€” HTTP 202)

{
  "success": true,
  "task_id": "task_abc123",
  "status": "processing",
  "message": "Task is still processing",
  "created_at": "2026-09-28T20:31:12.482Z",
  "started_at": "2026-09-28T20:31:40.101Z",
  "finished_at": null,
  "elapsed_seconds": 32.1
}

While the task is still waiting in the queue ("status": "pending"), the same response shape is returned with "started_at": null.

Response (Completed) Permalink to Response (Completed)

{
  "task_id": "task_abc123",
  "status": "completed",
  "success": true,
  "url": "https://example.com",
  "domain": "example.com",
  "title": "Example Domain",
  "content": "<!DOCTYPE html>...",
  "status_code": 200,
  "credits_used": 1.0,
  "credits_remaining": 99.0,
  "mode_used": "auto (request)",
  "created_at": "2026-09-28T20:31:12.482Z",
  "started_at": "2026-09-28T20:31:40.101Z",
  "finished_at": "2026-09-28T20:31:57.800Z",
  "elapsed_seconds": 45.3
}

Response (Failed) Permalink to Response (Failed)

{
  "task_id": "task_abc123",
  "status": "failed",
  "success": false,
  "error": "Connection timeout after 30 seconds",
  "credits_used": 0.5,
  "created_at": "2026-09-28T20:31:12.482Z",
  "started_at": "2026-09-28T20:31:40.101Z",
  "finished_at": "2026-09-28T20:31:51.900Z",
  "elapsed_seconds": 39.4
}
Use Cloud Storage otherwise task results are stored on our servers for a limited time: a task record is deleted once its result has been fetched (success or failure), and otherwise expires 10 minutes after the last status write β€” after that, polling returns 404. If not using cloud storage retrieve your results promptly after completion.

Task Timing Permalink to Task Timing

Every response that belongs to a task β€” the task-status GET endpoints and the task-creating POST routes β€” carries four timing fields. They let you separate queue wait from actual work:

  • started_at βˆ’ created_at = time the task spent pending in the queue
  • finished_at βˆ’ started_at = time the worker spent processing
Field Type Set when Notes
created_at string, ISO-8601 UTC YYYY-MM-DDTHH:MM:SS.sssZ (always milliseconds, Z suffix) The task was accepted by the API (it entered pending) Never changes afterwards
started_at same format, or null A worker picked the task up (first move to processing) null the whole time the task is queued
finished_at same format, or null The task reached success or failure null until it finishes
elapsed_seconds number with one decimal, or null always finished_at βˆ’ created_at once finished; now βˆ’ created_at while running β€” it grows between polls and freezes when the task finishes

null means “not yet / not known”. Tasks created before these fields existed return null for all four rather than a made-up value.

The same data as headers Permalink to The same data as headers

Every task response also carries the timing as headers β€” X-Task-Created-At, X-Task-Started-At, X-Task-Finished-At and X-Task-Elapsed-Seconds β€” with the same values and format. Null values are omitted (headers can’t be null), so a queued task simply has no X-Task-Started-At header.

When the response body isn’t JSON β€” delivery=raw returns the page, screenshot or PDF itself as the body β€” the headers are the only place the timing appears. And in the rare case where a JSON body already uses one of the four field names for its own data (for example a saved config’s own created_at), the timing fields are not injected into that body at all β€” read them from the X-Task-* headers instead. A response body never mixes its own timestamps with the task’s.

Where the fields appear Permalink to Where the fields appear

  • All task-status GET endpoints β€” scraper, crawl, map, search, config generate, schemes and agent (e.g. GET /api/v1/scraper/tasks/{task_id}) β€” on every response: pending, processing, the 202 “result is being finalized”, success, failure and unknown-status.
  • All task-creating POSTs β€” scraper /realtime, crawl, map, search, config generate, scheme create/update and agent /converse:
    • async 202 (task accepted): created_at set, started_at/finished_at null, elapsed_seconds β‰ˆ 0.0
    • sync 202 after the wait limit (“still processing”): same shape; started_at is null while the task is still queued
    • sync success / failure within the wait limit: full timing including finished_at
  • Not present: responses where no task was created (validation errors, 402 insufficient credits, 429 concurrency limit), 404 task not found, and agent /request (it is stateless and never creates a task).

Polling for Results Permalink to Polling for Results

Python Example Permalink to Python Example

import time
import requests

def wait_for_task(task_id, api_key, max_wait=120, poll_interval=2):
    """Poll task status until completion"""
    url = f"https://scrape.evomi.com/api/v1/scraper/tasks/{task_id}"
    headers = {"x-api-key": api_key}
    
    start_time = time.time()
    
    while time.time() - start_time < max_wait:
        response = requests.get(url, headers=headers)
        data = response.json()
        
        status = data.get("status")
        
        if status == "completed":
            return data
        elif status == "failed":
            raise Exception(f"Task failed: {data.get('error')}")
        
        # Still processing
        time.sleep(poll_interval)
    
    raise TimeoutError(f"Task {task_id} did not complete in {max_wait} seconds")

# Usage
result = wait_for_task("task_abc123", api_key)
print(result["content"])

JavaScript Example Permalink to JavaScript Example

async function waitForTask(taskId, apiKey, maxWait = 120000, pollInterval = 2000) {
  const url = `https://scrape.evomi.com/api/v1/scraper/tasks/${taskId}`;
  const headers = { 'x-api-key': apiKey };
  
  const startTime = Date.now();
  
  while (Date.now() - startTime < maxWait) {
    const response = await fetch(url, { headers });
    const data = await response.json();
    
    if (data.status === 'completed') {
      return data;
    } else if (data.status === 'failed') {
      throw new Error(`Task failed: ${data.error}`);
    }
    
    await new Promise(resolve => setTimeout(resolve, pollInterval));
  }
  
  throw new Error(`Task ${taskId} did not complete in time`);
}

// Usage
const result = await waitForTask('task_abc123', apiKey);
console.log(result.content);

Go Example Permalink to Go Example

func WaitForTask(taskID, apiKey string, maxWait time.Duration) (map[string]interface{}, error) {
    url := fmt.Sprintf("https://scrape.evomi.com/api/v1/scraper/tasks/%s?api_key=%s", taskID, apiKey)
    
    startTime := time.Now()
    pollInterval := 2 * time.Second
    
    for time.Since(startTime) < maxWait {
        resp, err := http.Get(url)
        if err != nil {
            return nil, err
        }
        
        var data map[string]interface{}
        json.NewDecoder(resp.Body).Decode(&data)
        resp.Body.Close()
        
        status := data["status"].(string)
        
        if status == "completed" {
            return data, nil
        } else if status == "failed" {
            return nil, fmt.Errorf("task failed: %v", data["error"])
        }
        
        time.Sleep(pollInterval)
    }
    
    return nil, fmt.Errorf("task did not complete in %v", maxWait)
}

Raw vs JSON Responses Permalink to Raw vs JSON Responses

When polling for results, the response format depends on the delivery parameter you set in your original request.

Raw Delivery (Default) Permalink to Raw Delivery (Default)

If you used delivery=raw (or didn’t specify delivery), the polled result returns the raw content directly:

curl "https://scrape.evomi.com/api/v1/scraper/tasks/task_abc123?api_key=YOUR_API_KEY"

Response:

Content-Type: text/html; charset=utf-8
X-Credits-Used: 1.0
X-Credits-Remaining: 99.0
X-Task-Created-At: 2026-09-28T20:31:12.482Z
X-Task-Started-At: 2026-09-28T20:31:14.102Z
X-Task-Finished-At: 2026-09-28T20:31:28.771Z
X-Task-Elapsed-Seconds: 16.3

<!DOCTYPE html>
<html>
...
With raw delivery, you get only the contentβ€”no metadata like title, status code, or credits used in the response body. Metadata is available only in response headers β€” including the task timing, which appears only as X-Task-* headers.

JSON Delivery Permalink to JSON Delivery

If you used delivery=json, the polled result returns a structured JSON response with metadata:

curl "https://scrape.evomi.com/api/v1/scraper/tasks/task_abc123?api_key=YOUR_API_KEY"

Response:

{
  "task_id": "task_abc123",
  "status": "completed",
  "success": true,
  "url": "https://example.com",
  "title": "Example Domain",
  "status_code": 200,
  "credits_used": 1.0,
  "created_at": "2026-09-28T20:31:12.482Z",
  "started_at": "2026-09-28T20:31:14.102Z",
  "finished_at": "2026-09-28T20:31:28.771Z",
  "elapsed_seconds": 16.3
}
With JSON delivery, you must set include_content=true in your original request to receive the scraped content. Without this parameter, only metadata is returnedβ€”no HTML, Markdown, or other content.

Using Delivery and Include_content Permalink to Using Delivery and Include_content

When using JSON delivery mode, you must set include_content=true to receive the scraped content in the response.

Without include_content Permalink to Without include_content

curl "https://scrape.evomi.com/api/v1/scraper/realtime?url=https://example.com&delivery=json&api_key=YOUR_API_KEY"

Response omits content:

{
  "success": true,
  "url": "https://example.com",
  "title": "Example Domain",
  "status_code": 200,
  "credits_used": 1.0,
  "hints": ["Content omitted by default. Set include_content=true to include it."]
}

With include_content Permalink to With include_content

curl "https://scrape.evomi.com/api/v1/scraper/realtime?url=https://example.com&delivery=json&include_content=true&api_key=YOUR_API_KEY"

Response includes content:

{
  "success": true,
  "url": "https://example.com",
  "title": "Example Domain",
  "content": "<!DOCTYPE html>...",
  "status_code": 200,
  "credits_used": 1.0
}
Important: When using delivery=json, always set include_content=true if you need the scraped content (HTML, Markdown, etc.) in the response. Without this parameter, only metadata is returned to save bandwidth.

Batch Processing with Async Permalink to Batch Processing with Async

For large batches, submit all tasks first, then poll for results:

import asyncio
import aiohttp

async def submit_and_wait(urls, api_key):
    async with aiohttp.ClientSession() as session:
        # Submit all tasks
        task_ids = []
        for url in urls:
            async with session.post(
                "https://scrape.evomi.com/api/v1/scraper/realtime",
                headers={"x-api-key": api_key},
                json={"url": url, "async": True}
            ) as resp:
                data = await resp.json()
                task_ids.append(data["task_id"])
        
        # Poll for all results
        results = []
        for task_id in task_ids:
            result = await poll_task(session, task_id, api_key)
            results.append(result)
        
        return results

async def poll_task(session, task_id, api_key, max_wait=120):
    url = f"https://scrape.evomi.com/api/v1/scraper/tasks/{task_id}"
    headers = {"x-api-key": api_key}
    
    start = asyncio.get_event_loop().time()
    
    while asyncio.get_event_loop().time() - start < max_wait:
        async with session.get(url, headers=headers) as resp:
            data = await resp.json()
            
            if data["status"] == "completed":
                return data
            elif data["status"] == "failed":
                raise Exception(f"Task failed: {data.get('error')}")
        
        await asyncio.sleep(2)
    
    raise TimeoutError(f"Task {task_id} timed out")

# Usage
urls = ["https://example1.com", "https://example2.com", "https://example3.com"]
results = asyncio.run(submit_and_wait(urls, api_key))

Webhook Alternative Permalink to Webhook Alternative

Instead of polling, you can use webhooks to receive notifications when tasks complete. See Webhooks for details.

{
  "url": "https://example.com",
  "async": true,
  "webhook": {
    "url": "https://your-server.com/webhook",
    "webhook_type": "custom",
    "events": ["completed", "failed"]
  }
}

Best Practices Permalink to Best Practices

  1. Use appropriate poll intervals β€” Poll every 2-3 seconds, not faster
  2. Set reasonable timeouts β€” Most tasks complete within 60 seconds
  3. Handle failures gracefully β€” Check the error field in failed responses
  4. Retrieve results promptly β€” Task records are deleted once their result has been fetched, and otherwise expire 10 minutes after the last status write
  5. Use webhooks for batches β€” Avoid polling hundreds of tasks simultaneously
  6. Stop on a terminal response β€” Once finished_at is set, elapsed_seconds stops growing and re-polling returns the same frozen value (until the record is deleted or expires)