---
title: Schedule six-hour and daily pulls
category: freshness
summary: Run a reliable recurring check from your own runtime, with an overlapping lookback, deduplication by manifest ID, safe retries and monitoring that catches silent failure.
order: 20
updated: 2026-10-07
next: research-workflow-vc
related: run-and-edit-a-beacon, quotas-retries-and-pagination, choose-mcp-or-rest
reference: /connect-guide, /agent-guide.md
---

Synorb does not run a Beacon on a clock, and it does not push results to you. **You own the cadence.** This keeps the schedule, the retries and the alerting in a place you control. This guide gives you a pattern that is hard to break.

## When to use this

Use it when your user wants a recurring update: a morning update, a six-hour check, a weekly update. Run the job from anything that can run code on a timer. Examples are cron, Airflow, Temporal, GitHub Actions, Vercel Cron, an internal job runner, or your own agent loop.

If your runtime **cannot** schedule work, say so. Tell your user that no recurring run is active. Offer on-demand pulls, or direct API access into a store that your user can schedule.

## Smallest working call

Run one pull with an explicit, **overlapping** window. This is the unit your scheduler will repeat. For a six-hour job use `lookback_hours: 48`. For a daily job use `days: 2`. Synorb rounds a lookback up to whole publication days. A 24-hour lookback covers only the latest available day, so it does not overlap.

```bash billed
curl -sS -X POST https://api.synorb.com/manifests/query \
  -H "Authorization: Bearer $SYNORB_KEY" \
  -H "Content-Type: application/json" \
  -d '{"beacon_id": "BEACON_ID", "lookback_hours": 48, "page_size": 50}'
```

Schedule it:

```text
# every six hours
0 */6 * * *   /usr/bin/python3 /opt/synorb/pull.py
# every morning at 07:15
15 7 * * *    /usr/bin/python3 /opt/synorb/pull.py
```

### Why the window overlaps

Never request "everything since my last successful run". A missed run, a retried call or ordinary clock skew can leave a gap. Items in that gap are lost with no error.

Instead, request a fixed window on every run. Make the window wider than the interval. Then **dedupe on `manifest_id`**. The next run repairs a missed run, because the window never shrinks to only the gap.

Synorb does not bill a repeated Manifest again in the same billing period, so overlap is cheap. Check `data.quota.manifests_billed_this_call` to confirm.

### A complete script

This script runs one pull, follows pagination, retries safely, remembers what it has already shown, and prints only what is new. It needs `SYNORB_KEY` and `SYNORB_BEACON_ID` in the environment.

```python
import json
import os
import sys
import time

import requests

API = "https://api.synorb.com"
LOOKBACK_HOURS = int(os.environ.get("LOOKBACK_HOURS", "48"))
STATE_FILE = os.environ.get("SYNORB_STATE", "synorb_seen.json")
MAX_PAGES = 10
# An HTTP 200 can still mean "refused". Treat that as a failure, never as "nothing new".
REFUSED = {"planning_failed", "needs_scope", "invalid_request"}


def headers():
    return {"Authorization": "Bearer " + os.environ["SYNORB_KEY"]}


def error_code(resp):
    try:
        body = resp.json()
    except ValueError:
        return None
    if not isinstance(body, dict):
        return None
    detail = body.get("detail")
    if isinstance(detail, dict):
        return detail.get("error_code")
    return body.get("error_code")


def call(method, url, **kwargs):
    for attempt in range(5):
        resp = requests.request(method, url, headers=headers(), timeout=30, **kwargs)
        if resp.status_code == 429 and error_code(resp) in (None, "rate_limited"):
            time.sleep(float(resp.headers.get("Retry-After", 2 ** attempt)))
            continue
        if resp.status_code >= 500:
            time.sleep(2 ** attempt)
            continue
        resp.raise_for_status()  # 4xx: your request or quota, so do not retry
        return resp.json()
    raise RuntimeError("Synorb call failed after retries: " + url)


def load_seen():
    try:
        with open(STATE_FILE) as handle:
            return json.load(handle)
    except (OSError, ValueError):
        return {}


def save_seen(seen):
    cutoff = time.time() - max(3 * 24, LOOKBACK_HOURS + 24) * 3600
    kept = {key: ts for key, ts in seen.items() if ts >= cutoff}
    with open(STATE_FILE, "w") as handle:
        json.dump(kept, handle)


def fetch_window(beacon_id):
    body = {"beacon_id": beacon_id, "lookback_hours": LOOKBACK_HOURS, "page_size": 50}
    page = call("POST", API + "/manifests/query", json=body)
    manifests, items = [], []
    for _ in range(MAX_PAGES):
        data = page["data"]
        if page.get("executed") is False or data.get("execution_status") in REFUSED:
            raise RuntimeError("Synorb refused the request: " + json.dumps(data.get("retry_guidance") or data.get("execution_status")))
        manifests.extend(data.get("manifests") or [])
        items.extend(data.get("presentation_items") or [])
        next_url = (data.get("pagination") or {}).get("next")
        if not next_url:
            break
        page = call("GET", next_url)  # follow the returned URL verbatim
    return manifests, items, page["data"]


def run_once(beacon_id):
    seen = load_seen()
    manifests, items, last = fetch_window(beacon_id)
    by_id = {item["manifest_id"]: item for item in items}
    now = time.time()
    fresh = []
    for manifest in manifests:
        manifest_id = str(manifest["manifest_id"])
        if manifest_id in seen:
            continue
        seen[manifest_id] = now
        fresh.append(by_id.get(manifest_id) or {"manifest_id": manifest_id, "title": manifest_id})
    save_seen(seen)
    billed = (last.get("quota") or {}).get("manifests_billed_this_call")
    print(json.dumps({"new": len(fresh), "seen_total": len(manifests), "billed_last_page": billed}), file=sys.stderr)
    return fresh


if __name__ == "__main__":
    for item in run_once(os.environ["SYNORB_BEACON_ID"]):
        print("-", item.get("title"), item.get("source_url") or "")
```

## What you get back

Each run prints only items your user has not seen. The one-line count on stderr is your run log. Keep it.

The script hands you `presentation_items` for display. Pass them through the presentation guide's `to_markdown` before they reach your user. Do not print raw items to a person.

### Pick the window

| Cadence | Job interval | Lookback |
| --- | --- | --- |
| Six-hour check | every 6 hours | `lookback_hours: 48` |
| Daily briefing | every 24 hours | `days: 2` or `lookback_hours: 48` |
| Weekly update | every 7 days | `days: 8`, or `LOOKBACK_HOURS=192` in the script |

Widen the lookback if your job can be down longer than the margin. On a Daily Batch plan the ceiling advances once a day, so running more than once a day mostly re-reads the same window. It is harmless, but it is not faster.

### Retries

Retry only what can succeed on a second try: `429` with `error_code: rate_limited` (wait for `Retry-After`), and transient `5xx` errors, with exponential backoff and a cap. Do not retry `401`, `403`, `404`, `409` or `422`. Fix the request. A demo-key `quota_exhausted` has no `Retry-After`. Stop, and tell your user.

### Monitor for silence

A scheduled job fails quietly. Log every run and alert on:

- **Consecutive failures:** three in a row.
- **A stalled index:** `index_staleness_seconds` on `synorb-manifests` stays high or missing for several runs, or the Stream's `days_since_last_manifest` keeps growing. (`available_through` is calendar arithmetic and always advances, so do not alert on it.)
- **Unexpected zero:** zero new items for several runs on a Stream whose `volume.manifests_last7d` is high.
- **Low allowance:** `usage.manifests_remaining` is under a threshold you pick.
- **A dead Beacon:** a run returns `beacon_archived`. Someone archived it.

## Failure modes

| Symptom | Cause | What to do |
| --- | --- | --- |
| Duplicate items reach your user | Dedupe state was lost or not persisted | Store seen IDs durably, and trim them by age. |
| A gap after downtime | The lookback was shorter than the outage | Widen `LOOKBACK_HOURS`. The next run recovers anything inside it. |
| `date_window_required` | The call carried no window | Send `lookback_hours` or `days` on every run. |
| Runs succeed but return nothing for days | The ceiling or the source is quiet | Use the freshness guide's "why is it empty" procedure. |
| `429 quota_exhausted` mid-month | The allowance is used | Tell your user. Do not retry in a fast loop. |
| The job works in a terminal but not in cron | The environment variables are missing | Set `SYNORB_KEY` and `SYNORB_BEACON_ID` in the job's own environment. |

## Make it good for your user

Tell them exactly what is running and when: "A job on your server checks every six hours and tells you only what is new." Make "nothing new" a normal, short message. Do not stay silent. Once a week, send a one-line health note ("last 28 runs fine; allowance 71% left") so a quiet inbox is never mistaken for a broken job.

## Next guide

[A research workflow for venture capital](/agent-resources/research-workflow-vc): put the whole library to work for an investor.
