Complete from-scratch rewrite of the Insta360 X5 dual-fisheye stitching
pipeline. Previous attempt (stitch.py, compare.py) moved to old/.
Architecture: single backward-mapping remap fusing equirect projection,
per-scanline rolling shutter correction (32-point gyro-interpolated),
MEI projection (xi=2.0, 13-coeff extended model from protobuf sidecar),
Lanczos4 sampling. Blending via longitude_preference * coverage_depth
with symmetric gain correction.
PSNR vs Insta360 Studio ground truth:
22.87 dB at 1920x960, 22.41 dB at 7680x3840 (with bilateral denoise)
Key findings:
- No-flow alpha blending works as well as DIS optical flow. DIS
cannot track the repetitive mesh pattern, and the principled
longitude*depth blend handles most parallax naturally
- Multi-video IMU_TO_CAM calibration avoids single-video overfitting
- Bilateral denoise (d=9, s=40) gives +0.43-0.70 dB matching GT noise
- Translation-aware projection confirmed (sign: protobuf t_extrinsic
is FROM lens TO center) but sub-pixel effect at typical distances
- Remaining gap: scene-dependent tone mapping + 18px mesh parallax
at 3m (requires neural flow to resolve)
Insta360 X5 Stitching Pipeline
Linux stitcher for raw Insta360 X5 footage. Reads .insv files (two H.265 fisheye streams, IMU samples, a protobuf calibration sidecar) and produces stabilized equirectangular stills or video. No Insta360 Studio required.
PSNR against Studio's own output: 22.5 to 22.9 dB at 7680×3840.
Insta360 Studio is closed source and Windows/macOS only. Reproducing its output on Linux meant working out the .insv container, the protobuf-encoded MEI calibration, the IMU axis convention, and the stitching and blending math. The result is a single-file pipeline of about 1,200 lines that matches the reference within a dB.
Files
x5_pipeline.py. The pipeline.PIPELINE.md. Architecture: the twelve stages, the MEI model, the blending math, the known limitations.x5_pipeline.md. Longer notes: container format, IMU axis calibration, rolling shutter, optical flow experiments.old/. First implementation, plusFINDINGS.mdwith the reverse-engineering notes the rewrite is built on. Seeold/README.md.
Architecture
Everything fuses into one backward remap per output pixel (following the pattern from the Insta360 SDK and Qualcomm's stabilization patent):
- Parse the
.insvinto two H.265 streams and an IMU track. - Parse the
.pbsidecar for MEI calibration (xi = 2.0, 13 distortion coefficients per lens, per-lens extrinsics). - Derive per-frame stabilization from IMU gravity.
- Derive per-scanline rolling-shutter rotations, 32 SLERP keyframes across a 21 ms readout.
- For each output pixel: ray, stabilize, transform into the lens frame, MEI-project, distort, sample.
- Blend on longitude preference times coverage depth. No hardcoded feather width.
- Symmetric per-channel gain across the seam.
- Optional DIS optical flow for close-range parallax. Optional bilateral denoise.
Full treatment in PIPELINE.md.
Install
Python 3.12+, with ffmpeg and ffprobe on PATH.
uv sync
# or
pip install -e .
Usage
# single frame, full resolution, with denoising
uv run python x5_pipeline.py input.insv -o output.jpg -w 7680 --denoise
# full video
uv run python x5_pipeline.py input.insv -o output.mp4 -w 3840 --video
# stabilization off (required on un-calibrated hardware, see below)
uv run python x5_pipeline.py input.insv --no-stab -o output.jpg
# PSNR against a Studio-rendered reference
uv run python x5_pipeline.py input.insv --gt studio_render.mp4 -o output.jpg
The .insv needs to sit inside the camera's default layout:
DCIM/Camera01/VID_xxx_00_001.insv
MISC/Camera01/VID_xxx_00_001.insv.pb
Limitations
IMU calibration is camera-specific. The IMU_TO_CAM rotation in x5_pipeline.py was solved via Wahba's method against ground-truth gravity on one X5 unit. Unit-to-unit PCB mounting variation will degrade stabilization on other cameras. Pass --no-stab, or re-solve against a Studio render from your own hardware.
Close-object parallax. Around 18 px of ghosting at the stitch line for objects under 3 m, a function of the 30 mm inter-lens baseline. DIS flow helps but does not match Insta360's learned ai_stitch_model_v2.ins on repetitive patterns like fence mesh or foliage.
Per-frame ffmpeg decode. Each frame spawns its own ffmpeg process, about 2 s of overhead. Piped batch decoding is the obvious next step for video throughput.
License
MIT. See LICENSE.