The EQ was right, the phone moved: debugging a 5 dB gap in room correction
What RoomEQ does
RoomEQ is automatic room correction for the Mac. It plays a sine sweep through the speakers, an iPhone records it in Safari and streams raw PCM to the Mac over a WebSocket, a solver fits up to ten peaking filters against a target curve, and a real-time engine applies that EQ to all system audio through BlackHole. Then it measures again, because an EQ you can't verify is just a guess.
That last step is where this story starts. On real hardware (a MacBook Air, a small F&D 2.1 speaker system and an iPhone 15), one auto-tune run reported this:
Without EQ: RMS error 4.20 dB (bass 5.32 dB)
Iteration 1: 5 filters, preamp +0.0 dB, predicted RMS 1.78 dB. Re-measuring ...
Iteration 1: measured RMS error 6.78 dB (bass 9.27 dB)
The solver predicted the EQ would bring the error from 4.2 dB down to 1.8. Measured through the live EQ, it was 6.8: worse than no EQ at all. An earlier run at a lower volume had landed within 0.8 dB of its prediction. Something was off by 5 dB, almost all of it in the bass.
Three suspects
There were three plausible explanations, and they point at completely different fixes:
- The audio path. The EQ isn't reaching the speakers as designed: a wrong sample rate, a coefficient bug, the test signal injected at the wrong point in the engine.
- The speaker. Many small 2.1 systems have a "dynamic bass" or bass-limiter circuit that changes their response with level. If it boosts bass when the input gets quieter, it partly undoes every cut the EQ makes. The bad run was about 10 dB louder than the good one, which made this my first theory.
- The measurement. The before and after measurements don't describe the same thing.
Arguing about which one it was would have taken longer than measuring it.
Four sweeps, one position
The trick is to stop the phone from moving. If every sweep is taken at the same spot, the room's contribution is identical in all of them and cancels out of every difference:
- a: EQ bypassed
- b: through the EQ
- c: EQ bypassed, 10 dB quieter
- d: EQ bypassed again
Each suspect then gets its own comparison:
measured_eq = b.db - a.db # the EQ as it reaches the speakers
level_shape = c.db + quieter_db - a.db # does the speaker change with level?
repeat = d.db - a.db # noise floor: the yardstick
eq_off, eq_rms, eq_max = _stats(measured_eq - designed, mask)
lin_off, lin_rms, lin_max = _stats(level_shape, mask)
rep_off, rep_rms, rep_max = _stats(repeat, mask)
tol = max(1.0, 2.5 * rep_rms)
designed is the filter response computed straight from the biquad coefficients, preamp included. The mask limits the comparison to 30 Hz–4 kHz, and skips anything more than 15 dB below the median level: the region below the speaker's roll-off and deep room nulls, where a little noise swings the dB value without meaning anything.
Before trusting it on hardware, I tested the check against a simulated speaker with an automatic gain control on its bass band. It flagged it clearly: the quieter sweep came back with up to 19 dB more bass. Then it ran on the real system, with the phone left on the chair for all four sweeps:
repeatability shape error 0.78 dB rms (the noise floor)
EQ as designed? shape error 0.85 dB rms
same shape 10 dB quieter? error 0.66 dB rms
PASS
The EQ reached the speakers within 0.85 dB of the design, about as close as two identical sweeps get to each other. The speaker sounded the same at both levels, so there was no dynamic bass circuit, and my first theory was wrong. Running auto-tune's own scoring on this data gave a measured 1.71 dB against a predicted 1.71 dB.
So the solver, the engine, the audio path, the speaker and the scoring were all correct. That left the measurement.
The room changes from seat to seat
The answer was in a number I already had. At this one spot, the error without EQ was 7.7 dB. Averaged over three spots 40 cm apart, it was 4.2 dB. In a small room, bass is a pattern of standing waves, and moving half a metre can move you from a peak into a dip.
Auto-tune measured "before" at three phone placements, solved, applied the EQ, and then asked the user to place the phone at three positions again for "after". Of course they weren't the same three spots. With bass this sensitive to position, the difference between two sets of placements was bigger than the effect of the EQ. The 6.8 dB wasn't the EQ failing; it was a comparison between two different rooms.
Paired rounds
The fix applies the same idea as Verify to every auto-tune round: at each position, measure with the EQ off and with it on, without moving the phone in between.
for p in range(positions):
rig.prepare_position(p, positions)
for label, flt, pre in (("EQ off", (), 0.0), ("EQ on ", eq, preamp_db)):
m = _sweep(rig, ts, p, flt, pre, analysis, log)
Now every round reports its improvement against its own "before", taken at the same spots. The phone can land somewhere different in every round, and the comparison stays fair. As a bonus, the EQ-off sweeps from every round are pooled, so each re-solve sees more positions than the last.
The pooled data also gives a per-frequency spread across positions, and the solver now uses it. Where listening positions disagree strongly, the average is not a reliable thing to EQ towards, so those frequencies count for less:
if spread is not None:
self.w = self.w / (1.0 + (np.asarray(spread) / cfg.spread_ref_db) ** 2)
Alongside that came a cap of −12 dB on stacked cuts. The single-position solve from Verify had piled three filters into a −19 dB notch at 80 Hz: right for that one spot, and wrong 40 cm to either side.
The result
The next run, with paired rounds:
Round 1: same positions, RMS error without EQ 9.01 dB -> with EQ 3.45 dB
(improvement +5.56 dB)
Round 2: same positions, RMS error without EQ 8.95 dB -> with EQ 3.59 dB
(improvement +5.36 dB)
Consistent rounds, and a re-solve prediction of 3.0 dB that landed within 0.6 dB of the measurement. It also showed a new problem: the bass started 20–27 dB above target, more than the EQ is allowed to cut. The subwoofer level was simply too high. RoomEQ now detects that case and says so ("turn the subwoofer down by roughly N dB on the speaker itself"), because a knob beats any EQ. After turning it down, auto-tune measured 4.6 dB → 2.6 dB at the same positions.
What I took from it
The tempting fix was to tune the solver until prediction and measurement agreed. That would have hidden the real problem, which was never in the solver. What found it was a measurement designed so that only one thing could differ between the two sides of each comparison. Verify still ships as a button on the dashboard, because "is my room correction actually working?" deserves a better answer than a predicted number.