← All Writing
September 29, 202610 min read

How I Pull Live SPX Options Chains From Schwab’s API

Every weekday morning, a small script on my home server snapshots the same-day SPX options chain and saves it, so I can compare what the market expected to happen with what actually did.

YieldA local database of same-day SPX option prices, greeks and implied volatility from the first two hours of trading, plus an end-of-day comparison of implied vs realized volatility
DifficultyIntermediate (REST API calls, parsing a nested JSON response, SQLite, cron, and a little options math)
Total Cook TimeThe collector took one session. Getting it to actually collect took longer, because of a time-zone bug I’ll get to

Ingredients

What this is, and what it isn’t

This is a data-collection project: a script that reads public market prices and writes them to a file. It places no orders and knows nothing about any account. I’m writing it up because the plumbing (endpoints, response shapes, the ways it fails) is the part that’s hard to find written down. This is not investment advice. Nothing here is a recommendation to buy, sell or trade anything, and the numbers below describe a small sample of past data, not what will happen next. If you’re making real decisions with real money, talk to a licensed professional.

The Question: Does the Market Overprice the Day?

Every option price has a forecast buried inside it. If you know an option’s price, you can work backward to the amount of movement the market is pricing in. That number is called implied volatility, or IV. After the day ends, you can measure how much the price actually moved. That’s realized volatility, or RV.

The gap between the two is one of the most-studied things in finance, and I wanted to see it with my own data rather than take it on faith. Same-day SPX options (called “0DTE,” for zero days to expiration) are a clean place to look: they expire at 4pm, so by the end of the afternoon you know exactly how the forecast turned out.

So the plan: snapshot the chain through the first two hours of trading (9:30 to 11:30am Eastern), then record what actually happened after the close, and compare.

Why only the morning? Because the question is what the market expected the day to look like, and that forecast is cleanest early, before most of the day’s move has happened. By the afternoon, same-day options have so little time left that their prices shrink to pennies, and a one-cent change in price can swing the implied-vol reading wildly. Two hours of snapshots is enough to see how the forecast settles after the open, without collecting hours of increasingly noisy numbers.

What the Data Shows: Implied vs Realized

After the close, a second scheduled run fills in the day: the close, the day’s high and low, and two versions of realized volatility.

A separate analysis script then goes further. Instead of trusting Schwab’s IV number, it solves for IV itself from the at-the-money option’s midpoint price, using a proper Black-Scholes solver, and computes realized vol with all three range-based estimators over the trailing 20 days. Two independent readings of the same thing are a good way to catch a bug in either one.

Across every day with a complete end-of-day record, near-the-money implied vol averaged 15.7% over the morning, close-to-close realized averaged 9.3%, and range-based realized averaged 7.0%. Implied came in above close-to-close realized on roughly six days out of seven. The market priced in more movement than showed up, which matches decades of published research. What I find more interesting is how much that gap moves around day to day, and how much the answer depends on which realized-vol measure you pick. A couple of months is a start, not a conclusion.

Implied (near the money)Realized, close-to-closeRealized, range-based
0%10%20%30%40%JulAugSepImpliedRange-basedClose-to-close
Implied (blue) sits above both realized measures on most days. Close-to-close (orange) is the noisy one: it swings from near zero to the high 20s, while range-based (teal) stays in a tighter band.

Subtract one from the other and the gap is easier to see. On the average day, implied vol came in 6.4 percentage points above close-to-close realized.

-100+10+20JulAugSepaverage gap +6.4 pts
Implied beat realized on roughly six days out of seven, by 6.4 points on average. Each bar is one day: implied minus close-to-close realized. Blue means the market priced in more movement than showed up; red means less. The red days are the ones where the market moved more than expected.

🔧 Developer section: options-math traps

That’s the finding. The rest of this post is how the data gets collected, and the ways the collector broke along the way.

The Three API Calls

The whole collector is three requests to Schwab’s Market Data API, each with the access token in the header:

The chain request is where the choices live. Here’s what mine asks for:

GET /marketdata/v1/chains
symbol $SPX
contractType CALL # or PUT, one side per run
strikeCount 30 # strikes centered on the current price
strategy SINGLE
includeUnderlyingQuote true
fromDate today
toDate today # same-day expiry only

Pinning fromDate and toDate to today keeps the response small. Without them you get every expiration Schwab lists, which for SPX is a lot of JSON on every run.

One design choice worth explaining: I only pull one side of the chain per day, calls or puts. The reason is that near the current price, a call and a put at the same strike imply almost exactly the same volatility. That’s a pricing relationship called put-call parity, and it means the second side would mostly be a duplicate of the first. So the script picks one: if SPX opens above yesterday’s close it logs calls, and if it opens below, puts. That halves the data without losing the implied-vol reading, which is the number this whole project is about. It’s a rule for what to record, not a rule for what to do.

🔧 Developer section: reading the chain response

Storing It: Three Tables in One File

Each run writes one row for the moment (SPX price, VIX, which side it logged, minutes since the open) and thirty rows for the strikes (bid, ask, midpoint, spread, IV, delta, gamma, theta, vega, volume, open interest, and how far each strike is from the current price). A third table keeps one row per day, which the end-of-day run fills in with the close and the realized-vol numbers.

After a couple of months of trading mornings, the whole thing still fits in a database file smaller than most phone photos. SQLite is plenty here. It’s one file, it needs no server, and a couple dozen writes a morning is nowhere near straining it.

0dte-monitor.log
11:15:02 Snapshot at 11:15:02 ET
11:15:04 SPX opened UP → logging CALLs | VIX=16.08
11:15:05 Stored snapshot: 30 strikes, ATM IV=17.3%

16:15:04 EOD backfill: RV=2.7% IV-RV=+10.6%

A normal morning: one line per step, one summary line per snapshot. The end-of-day line is the payoff: how much movement was priced in vs how much showed up.

The Shape of a Morning

Saving the chain through the morning, instead of once a day, shows something a single snapshot never could: near-the-money implied vol drifts down as the morning goes on. Averaged across every day collected so far, it starts around 17% right after the open and settles near 15% by late morning.

15%16%17%9:3510:0510:3511:0511:2517.3%15.1%
The market prices in the most movement right after the open, then relaxes. Average near-the-money implied vol by time of day, across every morning collected (9:35 to 11:25am Eastern; the 9:30 snapshot, taken at the opening bell, is left out).

What Breaks

The code is short. Almost everything I learned came from the ways it failed, so here they are in the order they cost me.

1. The time-zone bug that collected nothing

The script checks the clock and exits unless it’s between 9:30 and 11:30am Eastern. That’s a sensible guard. The problem was the schedule that launched it: I wrote the cron entry as “run from 9 to 11,” thinking in market hours. But the server runs on Pacific time. So cron fired from 9 to 11am Pacific, which is 12 to 2pm Eastern, and every single run looked at the clock, saw it was outside the window, and quietly exited.

No errors. No crashes. A clean log, and no data. The fix was one line (schedule it for 6 to 8am Pacific), plus a comment in the code explaining why the hours look wrong.

Two guards can cancel each other out

Each piece was right on its own: the script gated on Eastern time, and the schedule was written in market hours. Together they never overlapped. If a job has both a schedule and a time check inside it, make sure they agree on the time zone. And on day one, check that rows are actually showing up, not just that the log is clean.

2. Market holidays

The script knows about weekends. It doesn’t know about holidays. On Labor Day it ran on its normal schedule all day, found no same-day expiration because the market was closed, and filled the log with empty-chain warnings. Harmless, because it doesn’t write anything when the chain is empty. But the empty-chain path is the only reason it didn’t write garbage, and that’s luck, not design.

The fix is to ask whether the market is open before doing anything else. There are two good ways to do that:

I’d use Schwab’s endpoint as the main check and the library as a backup. Either way, the rule is the same: a closed market should be a clean, logged “skipping today,” not a day of warnings.

3. Random 400s on a request that worked on the last run

Every so often, the chain endpoint returns 400 Bad Request for a stretch of the morning, even though nothing about the request has changed and the quote endpoint is answering fine. Then it clears up on its own.

A 400 is supposed to mean your request is malformed, so this one is easy to misread. What matters is that the script treats a failed chain as “skip this snapshot” rather than “crash” or “write a partial row.” A day with 18 of 24 snapshots is still a usable day. A day with 18 good rows and six half-written ones is not.

4. Timeouts, and the 7-day wall

A handful of runs hit a read timeout (10 seconds for quotes, 15 for the chain). Same treatment: log it, skip it, try again on the next run. And if the token ever lapses, every call fails at step one, which is exactly why the token post exists: a collector like this is only as reliable as the token underneath it.

5. The “empty database”

A Python gotcha worth knowing: the SQLite library creates a new empty file if the one you ask for doesn’t exist. Check the wrong path and you’ll find an “empty database” that you just made by looking for it.

🔧 Developer section: failure handling, in one list

What went fast

What needed patience

The token post was about keeping a door open. This one is about what walks through it: a few thousand numbers a day, saved on a schedule, that slowly turn into an answer to a question I actually care about: the market prices in more movement than it delivers, most days, and now I can see by how much.

← Back to all writing