Skip to documentation
Browse documentation

Jobs and results

Handle pending jobs and truncated previews.

View raw

MCP execution waits for a result by default, for up to approximately 55 seconds. Slow work returns a pending job_id instead of holding the request indefinitely.

Pending work

Call upscrape_get_job_result with the returned job ID. Continue with bounded backoff until the tool reports a completed or failed state.

If the capability returns a top-level array, pass offset and limit to page through it. The response includes result_pagination with the total, returned count, and has_more flag.

Do not call upscrape_execute again merely because the first result was pending. The job ID is the durable state handle.

Result preview limit

MCP result previews are limited to 24 KiB of encoded structured content by default. If a result exceeds that size, the response is explicitly marked as truncated. The accompanying text representation and protocol envelope add bytes on the wire.

Read the complete stored result through the same MCP connection, including when signed in with OAuth:

{"job_id": "JOB_ID", "format": "json_chunk", "byte_offset": 0, "chunk_bytes": 8192}

Append each result_chunk.text verbatim and request the returned next_byte_offset until has_more is false. Offsets count UTF-8 bytes; always use the returned offset because encoding and the response budget can shorten a chunk. Verify total_bytes and sha256, then decode the concatenated JSON. This works for nested objects as well as arrays and returns raw stored output without the preview's normalization or temporary signed links. Array offset/limit do not apply in this mode. Retrieval does not run the capability again.

Integrations holding an API key with jobs:read and permission for the job's platform/capability can also use GET /jobs/JOB_ID. OAuth tokens are bound to MCP.

Charging

Execution charges the capability's published credit cost once on successful completion. Pending responses, result retrieval, transport retries, and failed jobs do not add another capability charge.

Cancellation and progress

The server currently selects the JSON response option rather than SSE. It does not advertise progress or disconnect cancellation; clients use the explicit job handle and may stop polling without cancelling the job.

Select only the data you need

Use collection_path to select a nested array with a JSON Pointer, then apply offset and limit. Optional fields retains named JSON Pointers from each selected object. Returned projection keys are the exact pointer strings:

{"job_id":"JOB_ID","collection_path":"/items","offset":0,"limit":5,"fields":["/title","/price"]}

A projected item looks like {"/title":"Example","/price":12}. Missing fields are omitted, never fabricated. Escape / as ~1 and ~ as ~0 in pointer segments. Invalid or non-array collection paths return an actionable error with the same job handle, so correcting retrieval never repeats the paid scrape. Pagination describes the stored result, not an additional upstream fetch. Selection parameters do not apply to raw json_chunk reads.

collected_at records job completion. result_complete: false indicates additional pages or preview truncation. Check selection metadata to know whether fields or a nested collection were selected. Large results should be processed with a client script when the task requires the whole dataset.

Retry and wait controls

wait_ms controls how long this call waits; timeout_ms controls the job deadline. A pending response includes poll_after_ms. Stopping polling does not cancel work. Use a unique operation_key on execute and reuse it after connection failure, reconnecting or OAuth refresh. Requests without this key are deduplicated only for the same JSON-RPC ID, credentials and arguments, and only for about 15 minutes: after that the same ID starts a new job, because clients restart their request IDs in every session. Failed terminal jobs remain the same job on replay; starting a fresh attempt requires a deliberate new operation key.

Operation-key deduplication lasts while the original job record is retained. A different API key or OAuth client has a separate key namespace. Disconnecting does not cancel accepted work.

A failed job remains terminal on replay. retryable describes a potentially transient failure; it does not reopen the job. A deliberate new attempt uses a new operation key and is subject to the normal credit ceiling.

An execute call that returns temporarily_unavailable with retryable: false created no job and charged nothing: the capability cannot be served right now. Stop, use another capability, or try again later; an immediate retry returns the same error.

Credit prices and plans: Upscrape pricing