diff --git a/docs/superpowers/plans/2026-05-13-body-mesh-perf-round2.md b/docs/superpowers/plans/2026-05-13-body-mesh-perf-round2.md new file mode 100644 index 0000000..fa134f5 --- /dev/null +++ b/docs/superpowers/plans/2026-05-13-body-mesh-perf-round2.md @@ -0,0 +1,275 @@ +# Body Mesh Pipeline β€” Performance Round 2 + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Apply the remaining high-ROI optimizations from the 2026-05-13 critic review of the Body Mesh pipeline (O5 Swift memcpy bulk + O6 normals off-MainActor) and harden the TCP protocol with a 1-byte version field for forward compatibility. Plan 1 already shipped the highest-ROI item (O1 numpy serializer). + +**Architecture:** Swift-only perf rework on the receiver side; one cross-process protocol change (Python sender + Swift receiver both touch the version byte). The Python sender's wire format gets a new 1-byte `version=1` after `MAGIC`; the Swift receiver tolerates it and rejects unknown versions. Mesh-normal compute moves from MainActor to a `Task.detached` block, with the result handed back to MainActor for the GPU upload. + +**Tech Stack:** Swift 6 + RealityKit + Accelerate (vDSP for vector ops), Python 3.11+ numpy. + +**Reference critic findings (verified 2026-05-13):** +- 🟠 O5 β€” `OSCServer.swift:98-113` decodes 10475 vertices via a per-vertex `load(fromByteOffset:)` loop. Gain: 2-5 ms β†’ ~0.05 ms by replacing with a `memcpy` over the contiguous float32 vertex run. +- 🟠 O6 β€” `MeshRenderer.swift:153-183` recomputes vertex normals on the MainActor every frame even though topology is static. Gain: 3-8 ms of MainActor time freed. +- 🟑 Protocol gap β€” no version field; future incompat is silent. Cheap to fix while we're already touching both ends. + +**Out of scope:** +- O8 (reduce inference resolution 896 β†’ 512) β€” qualitative tradeoff, requires visual eval. +- CoreML conversion of Multi-HMR β€” 20-40h effort. +- Accelerate-vDSP SIMD of the actual cross-product loop (the off-MainActor move alone yields the budget gain; vDSP rewrite is a separate optimization). + +--- + +### Task 1: Protocol β€” version byte (forward compat) + +**Files:** +- Modify: `data_only_viz/smplx_osc_sender.py` (`_serialize_persons` β€” write version byte after MAGIC) +- Modify: `launcher/AV-Live-Body/Sources/AVLiveBody/OSCServer.swift` (`decode` β€” read+validate version byte) + +**Context:** Today the frame layout is `[u32 length][MAGIC "SMPX"][i32 n_persons][persons...]`. Insert a `u8 version=1` immediately after `MAGIC`. The Swift receiver rejects unknown versions with an error log and drops the frame. + +Frame layout becomes: +``` +[u32 length BEβ†’LE][MAGIC "SMPX" 4B][u8 version=1][u8 reserved=0 Γ—3][i32 n_persons][persons...] +``` + +The 3 reserved bytes keep `n_persons` aligned to 4-byte boundary (good for any future `memcpy` decoder). + +- [ ] **Step 1: Update Python sender** + +In `smplx_osc_sender.py._serialize_persons`, find the line that writes `MAGIC + n_persons`. Replace with: + +```python +PROTO_VERSION = 1 + +# In _serialize_persons: +header = MAGIC + bytes([PROTO_VERSION, 0, 0, 0]) + struct.pack("&1 | tail -10 +cd /Users/electron/Documents/Projets/AV-Live/launcher/AV-Live-Body && swift build -c release 2>&1 | tail -5 +``` + +Expected: 25 passed + 1 pre-existing fail; Swift clean. + +- [ ] **Step 5: Commit** + +```bash +git add data_only_viz/smplx_osc_sender.py data_only_viz/tests/test_smplx_osc_sender_serialize.py launcher/AV-Live-Body/Sources/AVLiveBody/OSCServer.swift +git commit -m "feat(protocol): add SMPX version byte" +``` + +(40 chars β€” fits.) + +--- + +### Task 2: Swift receiver β€” memcpy bulk decode (O5) + +**Files:** +- Modify: `launcher/AV-Live-Body/Sources/AVLiveBody/OSCServer.swift` (`decode` β€” per-vertex `load(fromByteOffset:)` loop replaced with bulk copy) + +**Context:** Current decode loop iterates 10475 times, calling `load(fromByteOffset:)` 3Γ— per vertex. Each `load` involves bounds checks and arithmetic. Replace with a single bulk read into a `[SIMD3]` of size 10475, OR keep as `Data.subdata(in:)` + `withUnsafeBytes` block copy if `SIMD3` proves awkward in Swift. + +- [ ] **Step 1: Identify the current vertex-decode loop** + +```bash +grep -n "fromByteOffset\|10475\|vertices\|N_VERTS" launcher/AV-Live-Body/Sources/AVLiveBody/OSCServer.swift +``` + +Find the `for _ in 0..] = Array(repeating: .zero, count: vertCount) +vertData.withUnsafeBytes { (raw: UnsafeRawBufferPointer) in + let src = raw.bindMemory(to: Float.self).baseAddress! + vertices.withUnsafeMutableBytes { dst in + memcpy(dst.baseAddress!, src, vertBytes) + } +} +offset += vertBytes +``` + +(Adapt to the existing `SMPLXPersonData` constructor β€” the field is likely `vertices: [SIMD3]`.) + +The `memcpy` is sound because `SIMD3` is layout-compatible with three contiguous `Float`s on Apple Silicon (both are 16-byte aligned but tightly packed when stored in a Swift Array). + +**Important:** if `SIMD3` has padding (it MAY be 16-byte aligned, occupying 16 bytes per element instead of 12), the memcpy is WRONG. Verify with: + +```swift +print("MemoryLayout>.stride =", MemoryLayout>.stride) +print("MemoryLayout>.size =", MemoryLayout>.size) +``` + +If `stride > 12`, fallback to a `[Float]` of 31425 elements and do the SIMD3 reshape at the consumer (MeshRenderer). The plan's spec is "vertices are layout-compatible"; if reality differs, escalate `NEEDS_CONTEXT`. + +- [ ] **Step 3: Build + run** + +```bash +cd /Users/electron/Documents/Projets/AV-Live/launcher/AV-Live-Body && swift build -c release 2>&1 | tail -10 +``` + +Expected: clean. + +If you wrote a quick stride probe (Step 2 caveat), run the built binary briefly to confirm `stride == size == 12`. If not, escalate. + +- [ ] **Step 4: Commit** + +```bash +git add launcher/AV-Live-Body/Sources/AVLiveBody/OSCServer.swift +git commit -m "perf(av-live-body): memcpy bulk vertex decode" +``` + +(46 chars β€” fits.) + +--- + +### Task 3: MeshRenderer β€” normals off MainActor (O6) + +**Files:** +- Modify: `launcher/AV-Live-Body/Sources/AVLiveBody/MeshRenderer.swift` + +**Context:** `computeVertexNormals` runs synchronously on MainActor as part of `updateMeshVertices`. Move it to a `Task.detached` (or a dedicated serial DispatchQueue) so it doesn't block the main thread. The result (`[SIMD3]` of length 10475) is handed back to MainActor for the GPU upload. + +- [ ] **Step 1: Inspect current `updateMeshVertices` and `computeVertexNormals`** + +```bash +grep -n "computeVertexNormals\|updateMeshVertices\|@MainActor\|withUnsafeMutableBytes" launcher/AV-Live-Body/Sources/AVLiveBody/MeshRenderer.swift +``` + +Identify: +- Whether `computeVertexNormals` is a free function, static, or instance method. +- Whether it captures `self` (must NOT, to be safely off-MainActor). +- The size of its input/output buffers. + +- [ ] **Step 2: Refactor to make `computeVertexNormals` Sendable + actor-free** + +If `computeVertexNormals` takes `(vertices: [SIMD3], indices: [UInt32]) -> [SIMD3]` as input, it should already be pure. Mark it explicitly `static` or move to a free function (and verify it's `Sendable` by inspection β€” no global mutable state). + +If it currently reads `self.indices` (the static face topology), pass `indices` as an explicit parameter. + +- [ ] **Step 3: Dispatch off-MainActor in `updateMeshVertices`** + +Rewrite `updateMeshVertices` (or whichever method drives the per-frame update): + +```swift +@MainActor +func updateMeshVertices(_ persons: [SMPLXPersonData]) async { + for person in persons { + // Compute normals OFF main actor: + let normals = await Task.detached(priority: .userInitiated) { [verts = person.vertices, idx = self.indices] in + Self.computeVertexNormals(vertices: verts, indices: idx) + }.value + // Back on MainActor β€” upload to GPU: + applyToGPU(person.pid, vertices: person.vertices, normals: normals) + } +} +``` + +(The exact method names depend on what the file already has. Read first.) + +`computeVertexNormals` becomes: + +```swift +nonisolated static func computeVertexNormals( + vertices: [SIMD3], + indices: [UInt32] +) -> [SIMD3] { + // ... existing body, no self access ... +} +``` + +If the caller was synchronous (`func updateMeshVertices(_:)` not `async`), make it async OR introduce a `Task { @MainActor in ... }` wrapper at the call site. + +- [ ] **Step 4: Verify no MainActor stall in the new path** + +The whole point is to keep `await Task.detached` from blocking. If the call site is `Task { @MainActor in renderer.updateMeshVertices(persons) }` (from `OSCServer.parseFrames`), the inner `await` correctly suspends the MainActor task while the detached task runs on a background thread. Good. + +Confirm visually that the path is correct. If the existing code does an `async let` pattern or other concurrency primitive, adapt. + +- [ ] **Step 5: Build + smoke check** + +```bash +cd /Users/electron/Documents/Projets/AV-Live/launcher/AV-Live-Body && swift build -c release 2>&1 | tail -10 +``` + +Expected: clean build. + +If there are Swift 6 concurrency warnings about captures, the `[verts = person.vertices, idx = self.indices]` capture list is the fix. + +- [ ] **Step 6: Commit** + +```bash +git add launcher/AV-Live-Body/Sources/AVLiveBody/MeshRenderer.swift +git commit -m "perf(av-live-body): normals off MainActor" +``` + +(43 chars β€” fits.) + +--- + +## Self-Review + +**1. Spec coverage:** Task 1 = protocol version byte (forward compat + cheap because we're touching both ends). Task 2 = critic's O5 (Swift memcpy bulk). Task 3 = critic's O6 (normals off MainActor). All listed audit nits addressed; O8 / CoreML / vDSP rewrite explicitly out of scope. + +**2. Placeholder scan:** No TBD / TODO / "fill in". Every step has runnable code or a clear escalation path. The SIMD3 stride caveat in Task 2 is real and gives the implementer concrete actions for both branches (`stride == 12` β†’ memcpy; otherwise β†’ escalate). + +**3. Type consistency:** `PROTO_VERSION = 1` (uint8) consistent across Python sender and Swift receiver. `SMPLXPersonData.vertices` type is whatever the existing code uses ([SIMD3] expected); Task 3 doesn't change it. The header offsets in Task 1 update the P3.T1 test in lockstep. + +--- + +## Execution Handoff + +Plan complete and saved. Two execution options: +1. **Subagent-Driven (recommended)** β€” fresh subagent per task, two-stage review. +2. **Inline Execution** β€” batch with checkpoints. + +Same approach as Plans 1, 3, 2. diff --git a/launcher/AV-Live-Body/Sources/AVLiveBody/BodyView.swift b/launcher/AV-Live-Body/Sources/AVLiveBody/BodyView.swift index 268dce2..2c652cb 100644 --- a/launcher/AV-Live-Body/Sources/AVLiveBody/BodyView.swift +++ b/launcher/AV-Live-Body/Sources/AVLiveBody/BodyView.swift @@ -102,6 +102,7 @@ struct BodyView: NSViewRepresentable { renderer.applyMaterialSettings( metallic: settings.meshMetallic, roughness: Float(settings.meshRoughness)) + renderer.applyWireframeSetting(settings.showWireframe) for entity in renderer.personEntities.values { anchor.addChild(entity) } diff --git a/launcher/AV-Live-Body/Sources/AVLiveBody/MeshRenderer.swift b/launcher/AV-Live-Body/Sources/AVLiveBody/MeshRenderer.swift index 4b4c917..4ab9894 100644 --- a/launcher/AV-Live-Body/Sources/AVLiveBody/MeshRenderer.swift +++ b/launcher/AV-Live-Body/Sources/AVLiveBody/MeshRenderer.swift @@ -38,7 +38,21 @@ final class MeshRenderer: ObservableObject { arr[tri + 2] = tmp } self.faces = arr - NSLog("AV-Live-Body: loaded %d face indices (%d triangles)", n, n / 3) + // Pre-compute les indices d'aretes pour le wireframe : + // chaque triangle (a,b,c) -> 3 lignes (a,b)(b,c)(c,a). + var lines = [UInt32]() + lines.reserveCapacity(n * 2) + var i = 0 + while i + 2 < n { + let a = arr[i], b = arr[i + 1], c = arr[i + 2] + lines.append(a); lines.append(b) + lines.append(b); lines.append(c) + lines.append(c); lines.append(a) + i += 3 + } + self.wireframeIndices = lines + NSLog("AV-Live-Body: loaded %d face indices (%d triangles, %d wireframe edges)", + n, n / 3, lines.count / 2) } func startOSCServer() { @@ -56,6 +70,8 @@ final class MeshRenderer: ObservableObject { for (pid, _) in personEntities where !receivedPids.contains(pid) { personEntities.removeValue(forKey: pid) lowLevelMeshes.removeValue(forKey: pid) + wireframeMeshes.removeValue(forKey: pid) + wireframeEntities.removeValue(forKey: pid) } for p in persons { let entity: ModelEntity @@ -89,8 +105,12 @@ final class MeshRenderer: ObservableObject { } private var lowLevelMeshes: [Int: LowLevelMesh] = [:] + private var wireframeMeshes: [Int: LowLevelMesh] = [:] + private var wireframeEntities: [Int: ModelEntity] = [:] + private var wireframeIndices: [UInt32] = [] private var currentMetallic: Bool = false private var currentRoughness: Float = 0.6 + private var currentShowWireframe: Bool = false /// Pousse les nouveaux parametres de materiau a chaque entity vivant. /// Appelle par BodyView a chaque updateNSView pour permettre des @@ -106,6 +126,14 @@ final class MeshRenderer: ObservableObject { } } + /// Active/desactive le rendu fil de fer pour toutes les personnes. + func applyWireframeSetting(_ enabled: Bool) { + currentShowWireframe = enabled + for (_, wf) in wireframeEntities { + wf.isEnabled = enabled + } + } + private func makeEntity(pid: Int) -> ModelEntity { let material = SimpleMaterial(color: colorForPid(pid), roughness: .init( @@ -125,9 +153,63 @@ final class MeshRenderer: ObservableObject { materials: [material] ) } + // Cree l'entity wireframe sibling (cache jusqu'a toggle ON) + if let wfMesh = createWireframeMesh(vertices: initial) { + wireframeMeshes[pid] = wfMesh + if let wfResource = try? MeshResource(from: wfMesh) { + let wfMaterial = UnlitMaterial(color: NSColor.white) + let wfEntity = ModelEntity( + mesh: wfResource, materials: [wfMaterial]) + wfEntity.isEnabled = currentShowWireframe + entity.addChild(wfEntity) + wireframeEntities[pid] = wfEntity + } + } return entity } + /// Cree un LowLevelMesh en topologie .line avec les aretes de la + /// topologie SMPL-X. Vertex buffer initial = positions zero ; + /// updateMeshVertices se charge de la mise a jour live. + private func createWireframeMesh(vertices: [SIMD3]) -> LowLevelMesh? { + let posAttr = LowLevelMesh.Attribute( + semantic: .position, format: .float3, offset: 0) + let stride = MemoryLayout>.stride + let posLayout = LowLevelMesh.Layout( + bufferIndex: 0, bufferStride: stride) + let desc = LowLevelMesh.Descriptor( + vertexCapacity: vertices.count, + vertexAttributes: [posAttr], + vertexLayouts: [posLayout], + indexCapacity: wireframeIndices.count, + indexType: .uint32 + ) + guard let mesh = try? LowLevelMesh(descriptor: desc) else { + return nil + } + mesh.withUnsafeMutableBytes(bufferIndex: 0) { ptr in + let dst = ptr.bindMemory(to: SIMD3.self) + for (i, v) in vertices.enumerated() where i < dst.count { + dst[i] = v + } + } + mesh.withUnsafeMutableIndices { ptr in + let dst = ptr.bindMemory(to: UInt32.self) + for (i, idx) in wireframeIndices.enumerated() where i < dst.count { + dst[i] = idx + } + } + let bounds = BoundingBox(min: SIMD3(-2, -2, -2), + max: SIMD3(2, 2, 2)) + mesh.parts.replaceAll([ + .init(indexCount: wireframeIndices.count, + topology: .line, + materialIndex: 0, + bounds: bounds) + ]) + return mesh + } + private func fallbackMesh(vertices: [SIMD3]) -> MeshResource { var desc = MeshDescriptor(name: "smplx") desc.positions = MeshBuffer(vertices) @@ -151,6 +233,14 @@ final class MeshRenderer: ObservableObject { let n = min(dst.count, vertices.count) for i in 0...self) + let n = min(dst.count, vertices.count) + for i in 0..