Result Polling
When scraping tasks take longer than the timeout limit or when you submit async requests, you’ll need to poll for results. This guide covers how to retrieve results from background tasks.
When Polling Is Needed Permalink to When Polling Is Needed
Polling is required in these scenarios:
- Explicit async requests β You set
async=truein your request - Request timeouts β A synchronous request exceeds the timeout limit and auto-converts to async
Timeout Handling Permalink to Timeout Handling
Different modes have different timeout limits:
| Mode | Timeout | What Happens After |
|---|---|---|
request |
30 seconds | Auto-converts to async task |
browser |
45 seconds | Auto-converts to async task |
auto |
30-45 seconds | Depends on mode used |
Timeout Response Permalink to Timeout Response
When a synchronous request times out, you receive a 202 Accepted response:
{
"success": true,
"task_id": "task_abc123",
"status": "processing",
"message": "Task is taking longer than expected. Use task_id to check status.",
"check_url": "/api/v1/scraper/tasks/task_abc123",
"created_at": "2026-09-28T20:31:12.482Z",
"started_at": "2026-09-28T20:31:40.101Z",
"finished_at": null,
"elapsed_seconds": 45.3
}Submitting Async Requests Permalink to Submitting Async Requests
For long-running or batch jobs, submit async requests from the start:
Request Permalink to Request
curl -X POST "https://scrape.evomi.com/api/v1/scraper/realtime" \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com",
"async": true
}'Response Permalink to Response
{
"task_id": "task_abc123",
"status": "processing",
"check_url": "/api/v1/scraper/tasks/task_abc123",
"created_at": "2026-09-28T20:31:12.482Z",
"started_at": null,
"finished_at": null,
"elapsed_seconds": 0.1
}Checking Task Status Permalink to Checking Task Status
Request Permalink to Request
curl "https://scrape.evomi.com/api/v1/scraper/tasks/task_abc123?api_key=YOUR_API_KEY"Response (Processing β HTTP 202) Permalink to Response (Processing β HTTP 202)
{
"success": true,
"task_id": "task_abc123",
"status": "processing",
"message": "Task is still processing",
"created_at": "2026-09-28T20:31:12.482Z",
"started_at": "2026-09-28T20:31:40.101Z",
"finished_at": null,
"elapsed_seconds": 32.1
}While the task is still waiting in the queue ("status": "pending"), the same response shape is returned with "started_at": null.
Response (Completed) Permalink to Response (Completed)
{
"task_id": "task_abc123",
"status": "completed",
"success": true,
"url": "https://example.com",
"domain": "example.com",
"title": "Example Domain",
"content": "<!DOCTYPE html>...",
"status_code": 200,
"credits_used": 1.0,
"credits_remaining": 99.0,
"mode_used": "auto (request)",
"created_at": "2026-09-28T20:31:12.482Z",
"started_at": "2026-09-28T20:31:40.101Z",
"finished_at": "2026-09-28T20:31:57.800Z",
"elapsed_seconds": 45.3
}Response (Failed) Permalink to Response (Failed)
{
"task_id": "task_abc123",
"status": "failed",
"success": false,
"error": "Connection timeout after 30 seconds",
"credits_used": 0.5,
"created_at": "2026-09-28T20:31:12.482Z",
"started_at": "2026-09-28T20:31:40.101Z",
"finished_at": "2026-09-28T20:31:51.900Z",
"elapsed_seconds": 39.4
}404. If not using cloud storage retrieve your results promptly after completion.Task Timing Permalink to Task Timing
Every response that belongs to a task β the task-status GET endpoints and the task-creating POST routes β carries four timing fields. They let you separate queue wait from actual work:
started_at β created_at= time the task spent pending in the queuefinished_at β started_at= time the worker spent processing
| Field | Type | Set when | Notes |
|---|---|---|---|
created_at |
string, ISO-8601 UTC YYYY-MM-DDTHH:MM:SS.sssZ (always milliseconds, Z suffix) |
The task was accepted by the API (it entered pending) |
Never changes afterwards |
started_at |
same format, or null |
A worker picked the task up (first move to processing) |
null the whole time the task is queued |
finished_at |
same format, or null |
The task reached success or failure |
null until it finishes |
elapsed_seconds |
number with one decimal, or null |
always | finished_at β created_at once finished; now β created_at while running β it grows between polls and freezes when the task finishes |
null means “not yet / not known”. Tasks created before these fields existed return null for all four rather than a made-up value.
The same data as headers Permalink to The same data as headers
Every task response also carries the timing as headers β X-Task-Created-At, X-Task-Started-At, X-Task-Finished-At and X-Task-Elapsed-Seconds β with the same values and format. Null values are omitted (headers can’t be null), so a queued task simply has no X-Task-Started-At header.
When the response body isn’t JSON β delivery=raw returns the page, screenshot or PDF itself as the body β the headers are the only place the timing appears. And in the rare case where a JSON body already uses one of the four field names for its own data (for example a saved config’s own created_at), the timing fields are not injected into that body at all β read them from the X-Task-* headers instead. A response body never mixes its own timestamps with the task’s.
Where the fields appear Permalink to Where the fields appear
- All task-status GET endpoints β scraper, crawl, map, search, config generate, schemes and agent (e.g.
GET /api/v1/scraper/tasks/{task_id}) β on every response:pending,processing, the202“result is being finalized”, success, failure and unknown-status. - All task-creating POSTs β scraper
/realtime, crawl, map, search, config generate, scheme create/update and agent/converse:- async
202(task accepted):created_atset,started_at/finished_atnull,elapsed_secondsβ0.0 - sync
202after the wait limit (“still processing”): same shape;started_atisnullwhile the task is still queued - sync success / failure within the wait limit: full timing including
finished_at
- async
- Not present: responses where no task was created (validation errors,
402insufficient credits,429concurrency limit),404task not found, and agent/request(it is stateless and never creates a task).
Polling for Results Permalink to Polling for Results
Python Example Permalink to Python Example
import time
import requests
def wait_for_task(task_id, api_key, max_wait=120, poll_interval=2):
"""Poll task status until completion"""
url = f"https://scrape.evomi.com/api/v1/scraper/tasks/{task_id}"
headers = {"x-api-key": api_key}
start_time = time.time()
while time.time() - start_time < max_wait:
response = requests.get(url, headers=headers)
data = response.json()
status = data.get("status")
if status == "completed":
return data
elif status == "failed":
raise Exception(f"Task failed: {data.get('error')}")
# Still processing
time.sleep(poll_interval)
raise TimeoutError(f"Task {task_id} did not complete in {max_wait} seconds")
# Usage
result = wait_for_task("task_abc123", api_key)
print(result["content"])JavaScript Example Permalink to JavaScript Example
async function waitForTask(taskId, apiKey, maxWait = 120000, pollInterval = 2000) {
const url = `https://scrape.evomi.com/api/v1/scraper/tasks/${taskId}`;
const headers = { 'x-api-key': apiKey };
const startTime = Date.now();
while (Date.now() - startTime < maxWait) {
const response = await fetch(url, { headers });
const data = await response.json();
if (data.status === 'completed') {
return data;
} else if (data.status === 'failed') {
throw new Error(`Task failed: ${data.error}`);
}
await new Promise(resolve => setTimeout(resolve, pollInterval));
}
throw new Error(`Task ${taskId} did not complete in time`);
}
// Usage
const result = await waitForTask('task_abc123', apiKey);
console.log(result.content);Go Example Permalink to Go Example
func WaitForTask(taskID, apiKey string, maxWait time.Duration) (map[string]interface{}, error) {
url := fmt.Sprintf("https://scrape.evomi.com/api/v1/scraper/tasks/%s?api_key=%s", taskID, apiKey)
startTime := time.Now()
pollInterval := 2 * time.Second
for time.Since(startTime) < maxWait {
resp, err := http.Get(url)
if err != nil {
return nil, err
}
var data map[string]interface{}
json.NewDecoder(resp.Body).Decode(&data)
resp.Body.Close()
status := data["status"].(string)
if status == "completed" {
return data, nil
} else if status == "failed" {
return nil, fmt.Errorf("task failed: %v", data["error"])
}
time.Sleep(pollInterval)
}
return nil, fmt.Errorf("task did not complete in %v", maxWait)
}Raw vs JSON Responses Permalink to Raw vs JSON Responses
When polling for results, the response format depends on the delivery parameter you set in your original request.
Raw Delivery (Default) Permalink to Raw Delivery (Default)
If you used delivery=raw (or didn’t specify delivery), the polled result returns the raw content directly:
curl "https://scrape.evomi.com/api/v1/scraper/tasks/task_abc123?api_key=YOUR_API_KEY"Response:
Content-Type: text/html; charset=utf-8
X-Credits-Used: 1.0
X-Credits-Remaining: 99.0
X-Task-Created-At: 2026-09-28T20:31:12.482Z
X-Task-Started-At: 2026-09-28T20:31:14.102Z
X-Task-Finished-At: 2026-09-28T20:31:28.771Z
X-Task-Elapsed-Seconds: 16.3
<!DOCTYPE html>
<html>
...X-Task-* headers.JSON Delivery Permalink to JSON Delivery
If you used delivery=json, the polled result returns a structured JSON response with metadata:
curl "https://scrape.evomi.com/api/v1/scraper/tasks/task_abc123?api_key=YOUR_API_KEY"Response:
{
"task_id": "task_abc123",
"status": "completed",
"success": true,
"url": "https://example.com",
"title": "Example Domain",
"status_code": 200,
"credits_used": 1.0,
"created_at": "2026-09-28T20:31:12.482Z",
"started_at": "2026-09-28T20:31:14.102Z",
"finished_at": "2026-09-28T20:31:28.771Z",
"elapsed_seconds": 16.3
}include_content=true in your original request to receive the scraped content. Without this parameter, only metadata is returnedβno HTML, Markdown, or other content.Using Delivery and Include_content Permalink to Using Delivery and Include_content
When using JSON delivery mode, you must set include_content=true to receive the scraped content in the response.
Without include_content Permalink to Without include_content
curl "https://scrape.evomi.com/api/v1/scraper/realtime?url=https://example.com&delivery=json&api_key=YOUR_API_KEY"Response omits content:
{
"success": true,
"url": "https://example.com",
"title": "Example Domain",
"status_code": 200,
"credits_used": 1.0,
"hints": ["Content omitted by default. Set include_content=true to include it."]
}With include_content Permalink to With include_content
curl "https://scrape.evomi.com/api/v1/scraper/realtime?url=https://example.com&delivery=json&include_content=true&api_key=YOUR_API_KEY"Response includes content:
{
"success": true,
"url": "https://example.com",
"title": "Example Domain",
"content": "<!DOCTYPE html>...",
"status_code": 200,
"credits_used": 1.0
}delivery=json, always set include_content=true if you need the scraped content (HTML, Markdown, etc.) in the response. Without this parameter, only metadata is returned to save bandwidth.Batch Processing with Async Permalink to Batch Processing with Async
For large batches, submit all tasks first, then poll for results:
import asyncio
import aiohttp
async def submit_and_wait(urls, api_key):
async with aiohttp.ClientSession() as session:
# Submit all tasks
task_ids = []
for url in urls:
async with session.post(
"https://scrape.evomi.com/api/v1/scraper/realtime",
headers={"x-api-key": api_key},
json={"url": url, "async": True}
) as resp:
data = await resp.json()
task_ids.append(data["task_id"])
# Poll for all results
results = []
for task_id in task_ids:
result = await poll_task(session, task_id, api_key)
results.append(result)
return results
async def poll_task(session, task_id, api_key, max_wait=120):
url = f"https://scrape.evomi.com/api/v1/scraper/tasks/{task_id}"
headers = {"x-api-key": api_key}
start = asyncio.get_event_loop().time()
while asyncio.get_event_loop().time() - start < max_wait:
async with session.get(url, headers=headers) as resp:
data = await resp.json()
if data["status"] == "completed":
return data
elif data["status"] == "failed":
raise Exception(f"Task failed: {data.get('error')}")
await asyncio.sleep(2)
raise TimeoutError(f"Task {task_id} timed out")
# Usage
urls = ["https://example1.com", "https://example2.com", "https://example3.com"]
results = asyncio.run(submit_and_wait(urls, api_key))Webhook Alternative Permalink to Webhook Alternative
Instead of polling, you can use webhooks to receive notifications when tasks complete. See Webhooks for details.
{
"url": "https://example.com",
"async": true,
"webhook": {
"url": "https://your-server.com/webhook",
"webhook_type": "custom",
"events": ["completed", "failed"]
}
}Best Practices Permalink to Best Practices
- Use appropriate poll intervals β Poll every 2-3 seconds, not faster
- Set reasonable timeouts β Most tasks complete within 60 seconds
- Handle failures gracefully β Check the
errorfield in failed responses - Retrieve results promptly β Task records are deleted once their result has been fetched, and otherwise expire 10 minutes after the last status write
- Use webhooks for batches β Avoid polling hundreds of tasks simultaneously
- Stop on a terminal response β Once
finished_atis set,elapsed_secondsstops growing and re-polling returns the same frozen value (until the record is deleted or expires)