- 20 Jul, 2026 3 commits
-
-
Andrey Filippov authored
CLAUDE: item-4 R3 - RT LoG stream into the offline detector (det_rt_log_file) + det_ema_gate sizing stats New saved param curt.det_rt_log_file: the mode-3 detector consumes a pose-RT run's -CUAS-RT-LOG detection frames (LoG - avg) instead of its own subtract-avg+LoG; with the merged stack present and curt_subtract_avg on it still computes the offline frames and prints the RELATIVE-tol compare (ruling 3.3 - offline subtracts PRE-LoG, RT subtracts POST-LoG, equal by linearity except NaN edges) before swapping the RT frames in. No merged stack = the RT stack ingests as primary input. Outputs tagged -RTLOG; missing timestamps / synth_src combinations fail LOUD. det_log_save runs now also accumulate a |LoG-avg| log-histogram and print the post-LoG noise floor (sigma, p50..p99.99, max) + det_ema_gate candidates (4/5/6 sigma) and the percentile the current gate admits - the sizing data ruling 3.1 asked for (dial stays Andrey's). Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
R2 (rulings 3.1/3.3/3.4): CuasPoseRT arms the LoG+subtract-average stage with the render - taps from the static getLoGKenel (psf_radius/n_sigma, the same params detection uses); curt.det_avg_file (prior-run raw-domain average) LoG'd in Java DOUBLE at arm = the post-LoG base; gated EMA refresh (det_ema_rate 0.02 / det_ema_gate 10.0 placeholder dial, ungated = loud warning); detection frame (LoG - avg) stays GPU-resident. Post-loop saves -CUAS-RT-AVG (NaN-aware run average + count = the next run's prior file, skipped if the stage disabled mid-run) and the -CUAS-RT-LOG debug stack (det_log_save, ruling 3.4 test-only). New saved params det_log/det_log_save/ det_avg_file/det_avg_save/det_ema_rate/det_ema_gate (all 5 touch-points). Fold-in 1 (Session-54 close): the whole render ARM (renderSetup allocs + template + reference camera + R2 arm) hoisted PRE-LOOP next to renderProcPrepare - the scene-0 101/108 ms alloc-class outlier is gone; content-neutral (cv static per sequence; pXpYD_center + setReferenceGPU precede the arm). In-loop lazy arm removed (armed assertion, same fail-safe). Fold-in 2: GpuQuadJna.renderCameras static-table latch - first scene uploads radial/rByRDist once, then per-scene calls pass null and the native fast path sends meta+ERS pinned+async (renderSetup resets the latch per sequence). GpuQuad/GpuQuadJna/TpJna: renderLogSetup/renderLogGet surface. Stage0 EXPECTED_KERNELS 42 -> 45 (42 was stale - the render_dp bump to 43 was missed) and Stage0 passes 45/45. mvn package + test PASS; native gates in tile_processor_gpu 8840db8 (case log_pipe ALL PASS, sanitizer 0). Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
- GpuQuadJna.renderProcPrepare(spec): lazy second module (tp_create_module2, explicit SM-carve spec) + render-only TpProc bound to the main proc (tp_proc_render_bind_source); all render* entries route to it when present. Bound proc destroyed FIRST on close. Optional -Dtp.green.pose carves the MAIN module (default unset - INIT/pre-RT phases would be confined too). - CuasPoseRT: renderProcPrepare BEFORE the timed loop (NVRTC compile never lands in scene 0's slot); fail-safe = render stays on the main module. - New saved param curt.det_render_green (default "8+8" per the ratified MAP v2 pose-8/render-8/rest-20): empty/"0" = R1a single-module shape (A/B). - Native gates (tile_processor_gpu 0a0d509): carve_combo b0-b4 bind arrangements ALL PASS bit-exact incl. the PLAIN-pose + carved-render production mirror; full suite ALL PASS; sanitizer 0. - Step-4 gate (Andrey's run): 497 Done lines byte-identical to the 07/16 22:52 record with det_render ON + capacity re-measure. Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
- 19 Jul, 2026 3 commits
-
-
Andrey Filippov authored
The combined-image render (ROADMAP item 4, R1 shape confirmed by Andrey 07/19: pose keeps its own 1/60 slot; depth-2 pipeline, +1 frame detection latency) now runs from CuasPoseRT: after fitted[nscene] lands (fitted OR coasted - every scene renders, the no-skip rule), the hook collects the previous scene's render (normally already done - it ran during this scene's DP slot) and enqueues this scene's full-frame chain NON-BLOCKING on the low-priority render stream via GpuQuadJna.renderScene. - CuasRtParameters: curt.det_render (default ON; JNA lean path only, fail-safe disables the stage for the run) + curt.det_render_save (DEBUG, TEST-ONLY blocking D2H per ruling 3.4 -> -CUAS-RT-DETINPUT stack). - CuasPoseRT: per-sequence arm (renderSetup policy + FULL template = all tiles with valid disparity, the renderSceneVirtual selection=null convention + reference camera); per scene: render-private scene-camera snapshot (renderCameras - the pose loop's residents move on to N+1 before render(N) executes), shared uniform-MB descriptor (step 4 of poseMeasureSetup extracted verbatim into uniformMbDescriptor() - ops and order unchanged, both task streams carry the identical MB pair), ring-set captured before the D5 overlap flips it. NO per-scene prints - pose Done lines stay byte-identical; one-line stage summary after the loop. - GpuQuad/GpuQuadJna/TpJna: renderSetup/renderTemplate/renderCameras/ renderScene/renderStatus/renderGet (base = unsupported, JNA implements). mvn package + test PASS. Native side: tile_processor_gpu d2e0395/59884af/ fe69ecc (render_dp chain, case render_pipe bit-exact, sanitizer 0, run_cases ALL PASS, bench 4.30 ms/scene). Real-scene gate = R1c (Andrey's routine run): Done-lines byte-identical vs the 07/16 22:52 record + 60/s with the stage ON + -CUAS-RT-DETINPUT A/B vs CuasRender. Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
Real per-scene file I/O inside the measured loop (the honest RT test, and the production shape: 4 reader threads = the 4xSATA-2 device groups of 4 consecutive sensors). New CuasSceneFeeder parses the single-strip 16-bit TIFF layout ONCE from file #1 (strip@557, 656640 B, big-endian, telemetry row 0 - verified constant over the whole segment and bit-exact vs the production ImagejJp4Tiff reader), then preads raw strip bytes straight into the native pinned u16 staging (PosixIo/libc - the kernel's read IS the pinned write, half the pinned-write bytes of the float path). The GPU converts (byte-swap) + applies the borrowed per-scene Row/Col (prior scales snapshotted at creation, exact double chain) + two-map conditions in ONE kernel (condition_lwir_u16, native 84d7ab6) - bit-identical to the preload path by the u16_cond case, so the Done-line byte-identity gate carries over. Cold-read honesty: posix_fadvise(DONTNEED) after every file - reruns never replay the page cache (the production RAM-based dual-port SATA feeder always serves NEW frames). Cold bench: 4 threads 4.4/5.4/7.3 ms/scene p50/p95/max - hidden under the ~12 ms DP entry via the D5 overlap, which now awaits the in-flight read (kicked one scene ahead at collect-safe points, keeping the one-scene staging slack contract). Arrival moderation (Andrey 07/18): in paced (pose_pace) mode the feeder is armed with the same 60 Hz epoch grid - the DATA, not just the loop, arrives at 60 Hz; the overlap then fires only when the next read already landed (never parks inside the DP window), i.e. exactly when running behind. New saved checkbox curt.pose_feed16 (default ON) selects the feeder; OFF = the legacy float preload (A/B). CPU-oracle fallback (heap reads + the exact preload float chain) covers cond_gpu OFF and any native failure. Stage0 42/42; mvn package PASS. Co-authored-by:Claude Fable 5 <claude-fable-5@anthropic.com>
-
Andrey Filippov authored
At every poseSequencePass start, explicitly drop the per-sequence GPU task template and the DP seeding key, re-arm the scene-0 DP oracle, and return the device image-set ring to set 0. A back-to-back rerun in the same JVM now seeds identically to a fresh launch (restores the simple byte-identity gate; fresh starts were already bit-exact, reruns wobbled 0.0024 px from carried-over warm state - the previous run's last D5 overlap leaves an arbitrary active set while the new run's ring bookkeeping assumes set 0), and a (re)started production RT sequence never inherits stale resident state. Co-authored-by:Claude Fable 5 <claude-fable-5@anthropic.com>
-
- 18 Jul, 2026 4 commits
-
-
Andrey Filippov authored
The RT/CUAS path runs on the JNA native backend (libtileproc.so, offline nvcc 12.8); the legacy JCuda GPUTileProcessor is not needed there and its NVRTC+cuLinkAddData dies on Blackwell/sm_120 (jcuda 12.6 emits compute_120 PTX the driver JIT rejects -> CUDA_ERROR_INVALID_PTX; no jcuda 12.8 exists). - New maybeBuildGpuTileProcessor(): returns null (one-shot console note) when -Dtp.backend=jna, else builds GPUTileProcessor exactly as before. GpuQuad. create()/createRectilinear() already ignore a null gpuTileProcessor in JNA mode, and every GPU_TILE_PROCESSOR use is null-guarded or JCuda-only (ComboMatch orthomosaic already refuses JNA mode explicitly). - All 19 sites in Eyesis_Correction now call the guard. GUARD, not removal: jcuda mode (default, and the FOPEN/oracle path on its own branch) is byte-for-byte unchanged, so JCuda stays available as the FOPEN JCuda->JNA migration oracle. RT runs (-Dtp.backend=jna) no longer touch JCuda -> no INVALID_PTX regardless of menu path. mvn package + test PASS. Co-authored-by:
Claude Opus 4.8 <claude-opus-4-8@anthropic.com> Co-authored-by:
Claude Fable 5 <claude-fable-5@anthropic.com>
-
Andrey Filippov authored
Andrey's observation on the first perfd-held run: CPU boosting to 4.9 GHz, 100 C. Calibration (openssl all-core): the i9-11900K spikes 4.8 GHz @ 85 C, hits 100 C within seconds and throttles to a wobbling 3.2-3.5 GHz - the 'performance' boost bias is itself a nondeterminism source AND rides Tjmax. - measure profile: CPU pinned via sysfs scaling_min=max (default 3.2 GHz = the measured all-core throttle equilibrium, so the pin holds even in the worst phase and runs cooler; PERFD_CPU_LOCK_KHZ) + the existing GPU lock. PPD 'performance' hold now only backs the max profile. - HOLD accepts bounds-checked per-hold overrides (HOLD measure cpu=KHZ gpu=MHZ; perfctl passes them through) so lock-value calibration needs no daemon reconfiguration/sudo. - Holds now per-connection with values from the newest holder of the winning profile; STATUS extended. User-mode tested: override parsing/apply, bounds rejection, permission errors surfaced, auto-release. Root paths activate on reinstall. Co-authored-by:
Claude Opus 4.8 <claude-opus-4-8@anthropic.com> Co-authored-by:
Claude Fable 5 <claude-fable-5@anthropic.com>
-
Andrey Filippov authored
Design ruled by Andrey 07/17/2026 (lives in imagej-elphel: the capability follows the hardware the project runs on - 224/137; other programs use the same localhost socket/CLI): - perfd/elphel_perfd.py: root systemd service; line protocol on 127.0.0.1:48890 (PING/STATUS/HOLD measure|max/RELEASE); a hold is TIED TO THE CONNECTION (client crash = auto-release), refcounted, strongest profile wins. 'measure' = CPU performance (power-profiles-daemon hold, same call as the old cpu_perf_hold.py it supersedes) + GPU persistence + clocks LOCKED (default 2550 MHz; 5060 Ti burst boost measured ~2595 @ 23 W/49 C - the lock is about run-to-run +-1% reproducibility for per-item margin evaluation, not speed). 'max' = performance, unlocked. ExecStopPost=nvidia-smi -rgc so a dead daemon never leaves clocks locked. - perfd/install.sh (one-time sudo, idempotent) + perfctl (unprivileged CLI: 'perfctl hold measure -- <cmd>') + README. - Java: com.elphel.imagej.common.PerfDaemon (plain TCP, Java-8, zero deps; Hold implements AutoCloseable); Eyesis_Correction startup probe (one console line - uses the daemon when installed, prints the sudo install command when not); CuasPoseRT preloaded RT harness holds 'measure' for the timed run, releases after the RT summary. Tested end-to-end user-mode (daemon on a test port): protocol, Java client, CPU hold engaged, GPU lock correctly permission-refused as non-root, auto-release on disconnect. GPU lock path activates once installed under root. mvn package + test PASS. Co-authored-by:
Claude Opus 4.8 <claude-opus-4-8@anthropic.com> Co-authored-by:
Claude Fable 5 <claude-fable-5@anthropic.com>
-
Andrey Filippov authored
Co-authored-by:
Claude Opus 4.8 <claude-opus-4-8@anthropic.com> Co-authored-by:
Claude Fable 5 <claude-fable-5@anthropic.com>
-
- 17 Jul, 2026 8 commits
-
-
Andrey Filippov authored
Overlap-ring design ratified 07/17 (handoffs/2026-07-17_pose_d5_overlap_ring_design.md): - TpJna/GpuQuadJna/GpuQuad: bindings for the D5a image-set ring (selectImageSet/execConditioningAsync/imageSetReadyRecord) and the D5b launch/collect split (execPoseSceneDpLaunch/Collect); base/JCuda = unsupported (serial order + blocking entry retained). - CuasConditioning.conditionPreloadedSceneToGpuOverlap: same staged upload + same conditioning kernel into ring slot img_set on the copy stream (bytes identical by construction - the D5a slot-equivalence case); restores the previous set on any failure. - CuasPoseRT: dpSceneEntry takes an overlap Runnable - launch -> condition the NEXT scene -> collect (the ~5.6 ms/scene conditioning+H2D stage hides under the ~12.5 ms chain); the D4 oracle replay stays blocking; the loop builds the runnable (preloaded harness only, next non-null scene), skips the serial conditioning for a pre-conditioned scene, and disables the overlap for the rest of the run on any failure (fail-safe; the previous set is restored, records unaffected). - CuasRtParameters: saved checkbox pose_h2d_overlap (default ON, under pose_scene_dp) = the A/B knob. Gates: mvn clean package + mvn test PASS; Stage0 40/40 (no new kernels); native run_cases.sh ALL PASS (img_ring + scene_dp_split + all existing). REAL-SCENE GATE PENDING: routine saves-OFF run, 497 Done lines byte-identical to the 07/16 22:52 record, expect whole ~20.9 -> ~15.5 ms/scene, ~63-65 scenes/s (60 FPS PASS on the mean). Co-authored-by:
Claude Opus 4.8 <claude-opus-4-8@anthropic.com> Co-authored-by:
Claude Fable 5 <claude-fable-5@anthropic.com>
-
Andrey Filippov authored
Binds the D3-validated tp_proc_exec_pose_scene_dp in TpJna/GpuQuadJna (+ base GpuQuad stubs, non-throwing negative return = the FAIL-safe hook) and adds the one-entry-per-scene branch to the lean pose path: - new saved curt.pose_scene_dp (default ON, 'Pose scene as ONE GPU DP entry'): eligibility = frozen lean shape only (fixed cycles, 1-inner cap, freeze 0, 3-angle lma_use_R, uniform MB, JNA backend, no armed capture, no debug-save holders); everything else falls back to the host-driven (B3) loop automatically. - scene 0 (per template/sequence) always runs the host-driven loop - it seeds the resident per-sequence state through the production C1/C2 register path - then is REPLAYED through the DP entry as the one-shot oracle: per-cycle packed rows vs the captured resident-step results, the final anchor vs the fitted angles, and the last-cycle peak rows keyed by packed index (the B3 order-independence rule), all bit-compared. PASS arms DP from scene 1; FAIL (or any mid-run native failure) disables DP for the run and prints why. - DP scenes reconstruct the scene bookkeeping from the returned trace by the exact leanFitScene/IntersceneLma double rules (candidate = packed[16..18] widened, dATR in double, RMS = packed[19..22], rejected step == the legacy runLma -1 -> coast), fetch/unpack the resident last-cycle peaks exactly like leanMeasure, and print the same QC/cycles/Done-line fields - the real-scene gate is 497 Done lines byte-identical to the 07/16 22:52 saves-OFF record. - gpuTaskBuild steps 1-5 extracted verbatim into poseMeasureSetup() (shared with the DP entry); new 'scene DP entry' profile stage; Stage0 kernel count synced 36 -> 40 (the native-only D1-D3 sessions added the DP kernels but could not touch the Java constant). Gates: mvn clean package PASS; Stage0 PASS 40/40; native run_cases.sh ALL PASS (pose_corr @tol 0 + dp_cycles + measure_dp + scene_dp). Real-scene gate = Andrey's next routine run: expect the D4 oracle PASS line after scene 0, 'DP scene entry active' from scene 1, poses/QC byte-identical, post ~19.5 -> ~8-10 ms/scene. Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
The 07/16 22:09 B3 gate run FALSE-FAILed the one-shot oracle: the naive POSITIONAL compare ignored that correlate2D_inter assigns corr rows via atomicAdd - row ORDER is legitimately nondeterministic between runs (the rung-1 'order-independent corr compare by packed index' rule); all 497 poses stayed BYTE-IDENTICAL to the B4 record while the run was driven by batched cycles, proving the chain end-to-end. Fixes: - oracle compares per packed corr index (HashMap row lookup): index SETS must match and every index-matched 8-float peak row must be bit-exact; - REAL oracle FAIL now disables the batched path for the run (fail-safe: a divergence must never drive production); - B3_CHAIN profile stage attributed from the pre-gpuTaskBuild timestamp (was a fresh start AFTER the chain -> recorded ~0 and leaked ~2 ms/ cycle into 'leanMeasure other') and printed in the profile summary (printSummary never printed the new stage). mvn PASS. Oracle PASS line expected on the next routine run. Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
leanMeasure batched branch (JNA + pose_mb_uniform + resident build, no armed pose_corr capture, no debug-fetch holders): gpuTaskBuild runs b3CycleSetup (gc/cv + the 4 LPF/LoG consts + peak de-bias + sensor mask + per-scene descriptor - the same bytes the granular path re-uploaded every cycle) then ONE execPoseMeasureCycle call replaces the whole granular GPU sequence (task_update, geometry/convert x2, consolidate, inter-corr, normalize, peak). Per-cycle D2H stays corr indices + peaks (fetchCorr2DPeaks = D2H-only twin of execCorr2DPeaks). Granular path unchanged and remains the fallback + debug/capture path; new profile stage 'batched measure chain (B3)'. ONE-SHOT ORACLE (pose_lma_debug>=1, first eligible cycle): the batched chain runs FIRST (idempotent - every stage is a deterministic function of resident state + pose), captures indices+peaks, the granular chain re-runs the identical sequence, and the two are compared BIT-EXACT - EXPECT indices IDENTICAL + 0 peak mismatches. Batched production cycles start only after the oracle reports (or pose_lma_debug<1). ImageDtt.setInterCorrLpfs extracted from interCorrTDResident (shared). Requires tile_processor_gpu ee0d275. mvn PASS; run_cases.sh ALL PASS. Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
setBayerImages (both float[][] and double[][][] variants) now funnels through uploadBayer: each sensor image is written into the pinned native staging slot (one JNI region copy) and committed with tp_proc_set_image_staged - the async pinned->device DMA of sensor k overlaps the convert/write of sensor k+1 (per-sensor copy/compute pipelining, design B4). Double-buffered slots flip per upload (RT inter-scene overlap provision). Falls back one-time to the legacy pageable path on missing natives or any staged failure (legacy fences first, so a partial staged upload is always fully rewritten). ONE-SHOT ORACLE (first staged upload): reads every uploaded sensor back and counts bit-mismatches vs the Java array - EXPECT 0 EXACTLY, plus the enqueue-side upload ms. Requires tile_processor_gpu 11ccdc2 (libtileproc.so rebuild); older libs downgrade gracefully. mvn PASS; native run_cases.sh ALL PASS incl. new bayer_staged case. Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
Replace the per-scene one-point getMotionBlur call (~5-6 ms/scene, B0-attributed to IntersceneLma prepareLMA scaffolding + two ~THREADS_MAX thread-array spawns) with OpticalFlow.getMotionBlurPoint: the identical double chain (setupERS x2, getInterRotDeriveMatrices, getDPxSceneDParameters, same-order rate dot product) run serially at the representative point - bit-for-bit the same result at ~us cost. uniformMotionBlur broadcasts it unchanged; degenerate return still falls back to legacy getMotionBlur (per-tile), so non-uniform behavior is untouched. One-shot oracle at pose_lma_debug>=1 (replaces the closed B0 MB warm-repeat): times the fast path, evaluates legacy getMotionBlur at the SAME representative point, prints max|d| (expected 0 exactly). IntersceneLma.INFINITY_DISPARITY becomes the single source for the 0.01 infinity threshold. Gate (design doc B2): oracle max|d| = 0, real-scene poses within the B1 bounds, MB setup ~6 -> <0.5 ms/scene. Needs Andrey's run to close. Co-authored-by:
Claude <claude@elphel.com> Co-Authored-By:
Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
leanMeasure (uniform-MB lean path, JNA backend): the per-cycle CPU front end - transformToScenePxPyD (full 5120-tile grid), setInterTasksMotionBlur + sort, and the two task H2D uploads - is replaced by ONE gpuQuad.execPoseTaskUpdate() call: pose_task_update projects the per-sequence template on-device from the Java-tracked measure pose and rewrites both resident task slots in place; geometry + convert then run on ACTIVATED slots with no upload (ImageDtt.interCorrTDResident). Per cycle the only H2D left is the 12-float pose vector. Per scene: setupERS + camera-block upload (rung C2 registers, shared with prepare_resident) + skeleton slot re-upload (~120 KB, guards against other pipeline stages using the task slots between scenes). The per-scene uniform-MB 6-float descriptor replicates the exact setInterTasksMotionBlur double crank math; margin/projection failures become task=0 holes (missing peak -> conditioning abstains, D3 design). Fallbacks: JCuda backend, pose_mb_uniform off, NaN/non-uniform MB (degenerate uniformMotionBlur fallback) -> unchanged legacy CPU build. - GpuQuad/GpuQuadJna/TpJna: execPoseTaskUpdate + activateTaskSlot (base returns false; JNA marshals rung-C2 nullable groups). - IntersceneLmaFloat.buildTasks: serial float task-build oracle (the Java float clone of the kernel, worldFromPixel -> pixelFromWorld + descriptor/margin/hole packing). - One-shot oracle at pose_lma_debug>=1: GPU stream vs the legacy Java DOUBLE build (txy-matched field compare, D3 gate <=1e-5 px print) and vs the float-serial clone; once-per-program path-active note. - pose_corr export stays the explicit diagnostic exception: armed runs read the GPU-built pre-offset streams back and capture task words RAW (new iterTasksPre or511 overload - a |511-ed hole would diverge in a tol-0 replay). mvn -DskipTests clean package: BUILD SUCCESS. Real-scene gate (D3, ratified): poses < 0.003 px drift, aggregate RMS unchanged at print precision, same QC count, 497/497 - Andrey's run. CUDA side: tile_processor_gpu lwir16_2 @ ef36688 (kernel + JNA API + two-tier tests, all regressions PASS incl. pose_corr @tol 0). Design: internal handoffs/2026-07-16_3b_measure_chain_residency_design.md. Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
CLAUDE: 3-B rung B0 - one-shot attribution micro-measurements (grid transform, task build, MB setup, conditioning split) Design: attic/imagej-elphel-internal/handoffs/2026-07-16_3b_measure_chain_residency_design.md (all 4 rulings approved 07/16). B0 = attribute before B1 deletes the stages: - grid transform (pose_lma_debug>=1, once/program): warm re-runs of transformToScenePxPyD full-grid vs selection-only x multi vs single thread - splits thread-array churn / full-grid-for-150-tiles waste / irreducible per-tile ERS work. - task build: warm multi vs single-thread setInterTasksMotionBlur rebuild. - MB setup: warm repeat of uniformMotionBlur (sizes the B2 descriptor win). - conditioning (once/program, both direct and preloaded paths): prepare / cpu-condition / setBayerImages H2D / gpu-condition split - sizes B4's pinned+async overlap honestly (only the H2D share can hide). One-shot: extra runs land in 'leanMeasure other' for a single call, steady state unchanged. No behavior change; mvn -DskipTests package PASS. Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
- 16 Jul, 2026 17 commits
-
-
Andrey Filippov authored
C3b gate (log 191867): azimuth RMS 0.0454 (C2 bar 0.0349, +30%), roll 0.0415 (+54%), NOT-settled 228 (C2: 51). Better than the cycle-1 freeze (0.0518/0.0542/267) but still fails the pre-declared C2-level gate. CONCLUSION: per-cycle conditioning re-derivation carries convergence dynamics through ALL cycles (228 scenes not settled by cycle 4 - 'cycle 2 ~= converged' does not hold for half the sequence); not a noise source to freeze away for ~0.6 ms/scene. pose_freeze_cycle default 0 = per-cycle (restores C2 quality); parameter + bit-exact light-path mechanism stay for DP-era experiments (previous-scene conditioning candidate). Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
Andrey 07/16: better inter-scene accuracy may pay off - keep the opportunity. New curt.pose_freeze_cycle ('Pose freeze conditioning at cycle'): 0 = never freeze = per-cycle re-derivation (highest accuracy, pre-C3 behavior); 2 = default (freeze at the post-first-correction measurement); 1 = falsified. Runtime-selectable - A/B/revert experiments need no rebuild. Wired via IntersceneLma.setPoseFreezeCycles from leanFitScene. Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
C3 gate result: freezing at cycle 1 regressed azimuth RMS 0.0349->0.0518 px (+48%), roll 0.0270->0.0542 (+100%), NOT-settled 51->267. Mechanism: cycle-1 peaks are measured at the blurriest, pre-correction pose - freezing them locks in the WORST conditioning; per-cycle re-derivation carried convergence signal, not just noise. Andrey's observation: the first outer cycle absorbs nearly the whole prediction error (single near-GN inner step, lambda=1e-3), so the cycle-2 measurement is at an essentially converged pose. C3b: cycles 1-2 run FULL prepares; the SECOND result freezes; cycles 3+ run light. Context (Andrey): exit_change_atr QC threshold was test-era; the application bar is 0.1 px - C2-level quality remains the gate target, revert to per-cycle stays the fallback. Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
Design 3-A4i rung C3 (native tile_processor_gpu HEAD): the cycle-1 resident prepare result carries the FROZEN conditioning; later cycles of the same scene run a light prepare (fresh measured offsets + fx/J at the moving pose; weights/eigen/selection/pull/reg/pure_weight untouched). Rulings (Andrey 07/17): eigen freezes WITH the weights (same trust-in-tile nature); a frozen-selected tile with no fresh valid peak abstains (y=fx, zero residual); reg weights + pure_weight freeze from cycle 1. Semantic gain: every cycle optimizes the SAME objective - conditioning no longer wobbles with per-cycle measurement noise. Provider/GpuQuad/GpuQuadJna gain a light flag; the per-scene instance naturally scopes the freeze. Gates: mvn package+test PASS; Stage0 36/36; native suite incl. new LIGHT bit-exact test PASS. Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
Design 3-A4i rung C2 (native tile_processor_gpu HEAD): - IntersceneLma: per-SEQUENCE cache (reference Camera + flattened centers, keyed on the reference ErsCorrection; the virtual center is static by construction - zero pose, zero rates) with an explicit resetPoseSequenceCache() re-armed at sequence start; per-SCENE cache (scene Camera in the per-scene instance: ERS rates never change across cycles - setupERS depends only on rates/line_time, verified, and the 3 adjusted angles travel in pose_vectors). setupERS + Camera.capture now run once per level instead of 4x/scene; the provider gets null for resident groups. Provider/GpuQuad/GpuQuadJna signatures gain explicit numTiles. - Dead work deleted: the threaded setEigenTransform build at the prepareLMA head runs only on the legacy/capture/MB fall-through (the GPU assemble builds the transform from resident peaks on the production path). - CuasPoseRT: resetPoseSequenceCache() at testPoseSequence start. Expected: prep setup 5.7 -> ~1.5 ms cycle-1 / ~0.3-0.7 cycles 2-4, prepare ~22.9 -> ~3-5 ms/scene; per-cycle PCIe ~120KB -> ~60B + result. Gates: mvn package+test PASS; Stage0 36/36; full native suite PASS. Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
IntersceneLma is constructed per scene; the pre-existing one-shot flags are static but poseLmaPrepareOracleReported was instance-scoped, so the 13:03:32 run executed a legacy capture cycle + PREPARE compare on EVERY scene (497 bit-exact passes - strong accidental validation, but per-scene legacy cost: prepare 37.3 vs 22.9 ms/scene, capacity 6.67 vs 7.31 scenes/s). Now static. Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
prepareLMA received imp.debug_level only, so prepare_capture never armed in the debug1 config (pose_lma_debug=1, imp.debug_level<=0) and the C1 gate run went resident from cycle 0 without printing the bit-exact compare. Now the same boost the step oracle gets. Run was otherwise clean: 497/497, both markers, zero anomalies, drift <=0.0023 px, aggregate RMS/QC identical. Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
Design 3-A4i rung C1 (native tile_processor_gpu 7db7cf6): - IntersceneLma: PoseLmaPrepareProvider + residentPrepare() (mirrors the getFxDerivs head: ERS pokes + setupERS + camera/centers capture; per-cycle H2D = camera state + pose vector + centers + 6 policy floats). Production lean prepare skips setSamplesWeights / both fx passes / WJtJ+reg+normalize / y-build entirely; weights/y stay GPU-resident; initial RMS seeds from the first resident step's packed[19,20] (kills the runLma first-step re-linearization = C0's 'LMA CPU remainder'). Invalid/unavailable falls back loudly to the legacy Java prepare. - One-shot oracle (pose_lma_debug>=1, first prepare): the capture cycle runs FULLY legacy, then the resident prepare captures its buffers and the new IntersceneLmaFloat.prepareResidentOracle (serial float clone) must match BIT-EXACTLY ('resident CUDA vs Java-float PREPARE: ... mismatches=0'). - lmaStep: prepared-resident steps pass NULL weights/y/eigen (no per-step H2D); first-block linearization skipped when prepared. - CuasPoseRT: provider wiring (slot 0xff + CORR_NTILE_SHIFT = host policy), new marker 'resident CUDA prepareLMA active'; cycle_rms_meas moved after runLma reading getInitialRms() (same value on both paths). - Stage0 kernel count 33 -> 36 (3 new prepare kernels). Gates: mvn package + test PASS; Stage0 36/36; full native suite in tile_processor_gpu 7db7cf6 (direct prepare BIT-EXACT, sanitizer 0 errors). Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
Design 3-A4i rung C0: IntersceneLma gains a null-safe ProfileSink (pattern- consistent with the providers; zero overhead when unset). The production prepareLMA overload emits setup+samples-weights / fx-pass-1 / WJtJ+reg+ normalize / fx-pass-2 samples (LMA path only - MB-vector calls excluded; y+RMS tail = derived remainder). Two cross-cutting probes: the poseFxProvider JNA roundtrip wall (all lean fx sites, incl. runLma's first-step re-linearization) and the setupERS pair. CuasPoseRT routes the sink into RtPoseProfile (6 new stages, printSummary extended). Expected from one run of the same config: split prepare's ~16 ms/call into JNA-roundtrip vs Java-CPU, sizing C1 honestly. Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
Design 3-A4i rung B (native tile_processor_gpu 5f47326): for the lean three-angle shape the GPU decision drives the Java state machine: - resident-valid production path: accept/stop from packed[23..24], RMS from packed[19..22] (floats widened exactly), parameters_vector from the resident candidate; NO Java candidate projection, NO re-decision, NO rejection-restore projection (GPU retained current state, device-side set-index commit). Decision equivalence was proven over 1,988 steps with zero mismatches (session-43 gate). - invalid resident result (singular solve / non-finite) = REJECTED-STEP semantics: caller raises lambda and retries/exits; no Java-double fallback for the lean shape (fallbacks only where free). General shapes and the no-resident paths are unchanged. - one-shot oracle preserved: at pose_lma_debug>=1 the first step falls through the legacy path once and prints the existing preparation/candidate/ RMS-decision comparisons; production steps never compute the double side. - marker updated: 'resident CUDA float LMA decision AUTHORITATIVE'. Gates: mvn package PASS; mvn test PASS (no test sources); Stage0 33/33; full native suite in tile_processor_gpu 5f47326. Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
Design 3-A4i rung A (2026-07-16 handoff): the 00:25:48 gate profile's 'leanMeasure other' (26.1 ms/call, 43% of post) was unprofiled debug-save data motion (pose_corr_save/pose_img_save). Changes: - leanFitScene passes corr_pd_out/img_out only on the LAST outer cycle in fixed-cycle mode (legacy convergence mode keeps every-cycle so an early exit still fills the holders). - New DEBUG_FETCH profile stage wraps perSensorImagesFromTD, the no-MB debug re-convert, and fetchNormalizedPD, so debug data motion can never hide in 'leanMeasure other' again (it joins measureAccounted). Armed pose_corr capture unaffected (fetch keyed on isArmed() unchanged). Expected gate: poses/records byte-exact, hyperstacks identical, capacity ~3.9 -> ~5.5 scenes/s in the debug config. Co-Authored-By:Claude Fable 5 <noreply@anthropic.com>
-
Andrey Filippov authored
Co-authored-by:Codex <codex@elphel.com>
-
Andrey Filippov authored
Co-authored-by:Codex <codex@elphel.com>
-
Andrey Filippov authored
Co-authored-by:Codex <codex@elphel.com>
-
Andrey Filippov authored
Co-authored-by:Codex <codex@elphel.com>
-
Andrey Filippov authored
Co-authored-by:Codex <codex@elphel.com>
-
Andrey Filippov authored
Co-authored-by:Codex <codex@elphel.com>
-
- 15 Jul, 2026 5 commits
-
-
Andrey Filippov authored
Co-authored-by:Codex <codex@elphel.com>
-
Andrey Filippov authored
Co-authored-by:Codex <codex@elphel.com>
-
Andrey Filippov authored
Co-authored-by:Codex <codex@elphel.com>
-
Andrey Filippov authored
Co-authored-by:Codex <codex@elphel.com>
-
Andrey Filippov authored
Co-authored-by:Codex <codex@elphel.com>
-