Complete from-scratch rewrite of the Insta360 X5 dual-fisheye stitching
pipeline. Previous attempt (stitch.py, compare.py) moved to old/.
Architecture: single backward-mapping remap fusing equirect projection,
per-scanline rolling shutter correction (32-point gyro-interpolated),
MEI projection (xi=2.0, 13-coeff extended model from protobuf sidecar),
Lanczos4 sampling. Blending via longitude_preference * coverage_depth
with symmetric gain correction.
PSNR vs Insta360 Studio ground truth:
22.87 dB at 1920x960, 22.41 dB at 7680x3840 (with bilateral denoise)
Key findings:
- No-flow alpha blending works as well as DIS optical flow. DIS
cannot track the repetitive mesh pattern, and the principled
longitude*depth blend handles most parallax naturally
- Multi-video IMU_TO_CAM calibration avoids single-video overfitting
- Bilateral denoise (d=9, s=40) gives +0.43-0.70 dB matching GT noise
- Translation-aware projection confirmed (sign: protobuf t_extrinsic
is FROM lens TO center) but sub-pixel effect at typical distances
- Remaining gap: scene-dependent tone mapping + 18px mesh parallax
at 3m (requires neural flow to resolve)
3.3 KiB
Insta360 X5 Stitching Pipeline
Linux stitcher for raw Insta360 X5 footage. Reads .insv files (two H.265 fisheye streams, IMU samples, a protobuf calibration sidecar) and produces stabilized equirectangular stills or video. No Insta360 Studio required.
PSNR against Studio's own output: 22.5 to 22.9 dB at 7680×3840.
Insta360 Studio is closed source and Windows/macOS only. Reproducing its output on Linux meant working out the .insv container, the protobuf-encoded MEI calibration, the IMU axis convention, and the stitching and blending math. The result is a single-file pipeline of about 1,200 lines that matches the reference within a dB.
Files
x5_pipeline.py. The pipeline.PIPELINE.md. Architecture: the twelve stages, the MEI model, the blending math, the known limitations.x5_pipeline.md. Longer notes: container format, IMU axis calibration, rolling shutter, optical flow experiments.old/. First implementation, plusFINDINGS.mdwith the reverse-engineering notes the rewrite is built on. Seeold/README.md.
Architecture
Everything fuses into one backward remap per output pixel (following the pattern from the Insta360 SDK and Qualcomm's stabilization patent):
- Parse the
.insvinto two H.265 streams and an IMU track. - Parse the
.pbsidecar for MEI calibration (xi = 2.0, 13 distortion coefficients per lens, per-lens extrinsics). - Derive per-frame stabilization from IMU gravity.
- Derive per-scanline rolling-shutter rotations, 32 SLERP keyframes across a 21 ms readout.
- For each output pixel: ray, stabilize, transform into the lens frame, MEI-project, distort, sample.
- Blend on longitude preference times coverage depth. No hardcoded feather width.
- Symmetric per-channel gain across the seam.
- Optional DIS optical flow for close-range parallax. Optional bilateral denoise.
Full treatment in PIPELINE.md.
Install
Python 3.12+, with ffmpeg and ffprobe on PATH.
uv sync
# or
pip install -e .
Usage
# single frame, full resolution, with denoising
uv run python x5_pipeline.py input.insv -o output.jpg -w 7680 --denoise
# full video
uv run python x5_pipeline.py input.insv -o output.mp4 -w 3840 --video
# stabilization off (required on un-calibrated hardware, see below)
uv run python x5_pipeline.py input.insv --no-stab -o output.jpg
# PSNR against a Studio-rendered reference
uv run python x5_pipeline.py input.insv --gt studio_render.mp4 -o output.jpg
The .insv needs to sit inside the camera's default layout:
DCIM/Camera01/VID_xxx_00_001.insv
MISC/Camera01/VID_xxx_00_001.insv.pb
Limitations
IMU calibration is camera-specific. The IMU_TO_CAM rotation in x5_pipeline.py was solved via Wahba's method against ground-truth gravity on one X5 unit. Unit-to-unit PCB mounting variation will degrade stabilization on other cameras. Pass --no-stab, or re-solve against a Studio render from your own hardware.
Close-object parallax. Around 18 px of ghosting at the stitch line for objects under 3 m, a function of the 30 mm inter-lens baseline. DIS flow helps but does not match Insta360's learned ai_stitch_model_v2.ins on repetitive patterns like fence mesh or foliage.
Per-frame ffmpeg decode. Each frame spawns its own ffmpeg process, about 2 s of overhead. Piped batch decoding is the obvious next step for video throughput.
License
MIT. See LICENSE.