Perception Method Library Overview
This directory is the method-level perception library. Each page should represent one technique, method, benchmark, or dataset-backed evaluation primitive. Broad synthesis pages in 30-autonomy-stack/perception/overview/ remain useful for system design, but this library is where individual methods get enough space for architecture, data, benchmarks, failure modes, deployment fit, Domain Fit, transfer notes for explicitly scoped ODDs, and sources.
Priority Ratings
Priority ratings are editorial reading and deployment triage signals. Learning answers what to read early for general autonomy understanding. Deployment answers what to evaluate early for AV deployment in the tagged context; it is not a certification, product-readiness, or all-domain average claim. If a method's deployment score is driven by a specific domain or stack role, the reason text should name that context.
| Method | Rating | Stage | Maturity | Reason |
|---|---|---|---|---|
| MinkowskiNet | Learning: ★★★★★ Deployment: ★★★★★ | classic-baseline | fielded-pattern | MinkowskiNet sparse-convolution U-Nets are the deployed industry-baseline backbone for 3D semantic segmentation, including offline aggregated-map labeling. |
| Availability-Aware Sensor Fusion | Learning: ★★★★☆ Deployment: ★★★★★ | deployment-pattern | prototype | Directly targets sensor degradation and availability-aware fusion. |
| LiDAR-MOS | Learning: ★★★★☆ Deployment: ★★★★★ | deployment-pattern | prototype | Moving-object segmentation is central to map hygiene and dynamic-scene handling. |
| KPConv | Learning: ★★★★★ Deployment: ★★★★☆ | classic-baseline | fielded-pattern | KPConv is the point-based-convolution workhorse for large-scale MLS/ALS point cloud semantic segmentation and a proven baseline for aggregated-map labeling. |
| Point Transformer V3 | Learning: ★★★★★ Deployment: ★★★★☆ | modern-core | prototype | Point Transformer V3 is the leading serialized-attention 3D backbone for large-scale point cloud semantic segmentation, including offline aggregated-map labeling. |
| RandLA-Net | Learning: ★★★★★ Deployment: ★★★★☆ | classic-baseline | fielded-pattern | RandLA-Net is the efficient large-scale point-cloud segmentation baseline, designed for million-point clouds and used across MLS survey workflows. |
| 2DPASS | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | 2DPASS is the reference training scheme for image-assisted LiDAR semantic segmentation with LiDAR-only inference. |
| 4DSegStreamer | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | 4DSegStreamer is rated for motion segmentation, scene flow, or dynamic-object perception workflows. |
| AutoOcc | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | AutoOcc is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks. |
| BEVDepth | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | Important depth-aware BEV bridge for camera-only 3D perception. |
| BEVDet | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | Baseline camera BEV detector that organizes many later BEV methods. |
| BEVStereo | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | BEVStereo is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks. |
| Cam4DOcc | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | Cam4DOcc is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks. |
| Conformal Boxes | Learning: ★★★★☆ Deployment: ★★★★☆ | deployment-pattern | prototype | Practical uncertainty wrapper for detection risk and release gates. |
| Cross-Domain LiDAR Scene Flow | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | Cross-Domain LiDAR Scene Flow is rated for motion segmentation, scene flow, or dynamic-object perception workflows. |
| Cylinder3D | Learning: ★★★★☆ Deployment: ★★★★☆ | classic-baseline | prototype | Cylinder3D is the cylindrical sparse-voxel baseline for single-scan spinning-LiDAR semantic segmentation. |
| DepthOcc | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | DepthOcc is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks. |
| DiffusionDrive | Learning: ★★★★☆ Deployment: ★★★★☆ | frontier | prototype | Truncated diffusion policy that makes diffusion-based end-to-end driving fast enough for real-time use. |
| Dynamic Occupancy Freespace | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | Dynamic Occupancy Freespace is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks. |
| FlashOcc | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | FlashOcc is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks. |
| FRNet | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | FRNet is the frustum-range network that matches RangeFormer's accuracy at roughly 5x the speed — the current real-time range-image SOTA. |
| GaussianFlowOcc | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | GaussianFlowOcc is rated for sparse, weakly supervised semantic occupancy with explicit temporal Gaussian flow. |
| GaussianOcc | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | GaussianOcc is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks. |
| GaussRender | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | GaussRender is rated for improving 3D occupancy geometry through efficient projective Gaussian-rendering supervision. |
| GraphBEV | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | GraphBEV is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks. |
| InsMOS | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | InsMOS is rated for motion segmentation, scene flow, or dynamic-object perception workflows. |
| Instantaneous Motion Perception | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | Instantaneous Motion Perception is rated for motion segmentation, scene flow, or dynamic-object perception workflows. |
| LiDAR-Camera Occupancy Fusion | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | LiDAR-Camera Occupancy Fusion is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks. |
| M2-Occ | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | M2-Occ is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks. |
| MambaMOS | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | MambaMOS is rated for motion segmentation, scene flow, or dynamic-object perception workflows. |
| Mask4D | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | Mask4D is rated for motion segmentation, scene flow, or dynamic-object perception workflows. |
| MotionSeg3D | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | MotionSeg3D is rated for motion segmentation, scene flow, or dynamic-object perception workflows. |
| Neural Scene Flow Priors | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | Neural Scene Flow Priors is rated for motion segmentation, scene flow, or dynamic-object perception workflows. |
| Open-Vocabulary Panoptic Occupancy | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | Open-Vocabulary Panoptic Occupancy is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks. |
| OpenAD | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | fielded-pattern | Open-world benchmark for corner cases and unseen categories. |
| RadarPillars | Learning: ★★★★☆ Deployment: ★★★★☆ | classic-baseline | prototype | Core radar-native detection baseline for weather-robust perception. |
| RenderOcc | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | RenderOcc is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks. |
| SalsaNext | Learning: ★★★★☆ Deployment: ★★★★☆ | classic-baseline | pilot-proven | SalsaNext is the canonical real-time range-image LiDAR semantic segmenter — the fast, embedded-friendly projection baseline. |
| SegNet4D | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | SegNet4D is rated for motion segmentation, scene flow, or dynamic-object perception workflows. |
| SelfOcc | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | SelfOcc is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks. |
| SOLOFusion | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | SOLOFusion is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks. |
| SparseBEV | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | SparseBEV is rated for efficient sparse-query multi-camera 3D detection where dense BEV memory is costly. |
| SparseDrive | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | Fully sparse end-to-end driving stack that unifies detection, tracking, online mapping, and motion planning. |
| SparseOcc | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | SparseOcc is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks. |
| Spatiotemporal Memory Occupancy Flow | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | Spatiotemporal Memory Occupancy Flow is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks. |
| SPVCNN | Learning: ★★★★☆ Deployment: ★★★★☆ | classic-baseline | prototype | SPVCNN's sparse point-voxel convolution is the efficient point-voxel-hybrid baseline for real-time LiDAR semantic segmentation. |
| Streaming Gaussian Occupancy | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | Streaming Gaussian Occupancy is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks. |
| StreamingFlow | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | StreamingFlow is rated for motion segmentation, scene flow, or dynamic-object perception workflows. |
| StreamMOS | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | StreamMOS is rated for motion segmentation, scene flow, or dynamic-object perception workflows. |
| Superpoint Transformer | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | Superpoint Transformer is rated for large-scale 3D semantic segmentation of registered point clouds where whole-scene context and very small models matter. |
| SurroundOcc | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | Foundational camera occupancy reference for planning-facing perception. |
| TEOcc | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | TEOcc is rated for radar-camera semantic occupancy and temporal robustness under degraded perception conditions. |
| TPVFormer | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | TPVFormer is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks. |
| TrackOcc | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | TrackOcc is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks. |
| WaffleIron | Learning: ★★★★☆ Deployment: ★★★★☆ | modern-core | prototype | WaffleIron is a deliberately simple LiDAR segmentation backbone built from standard dense 2D convolutions, easy to implement and deploy. |
| 3D-KNN Blind-Spot Desnowing | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | pilot-proven | 3D-KNN Blind-Spot Desnowing is rated for cleaning, stress testing, or failure detection in degraded perception conditions. |
| 4D Radar Road Boundaries and Freespace | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | prototype | 4D Radar Road Boundaries and Freespace is rated for alternative-sensor perception and adverse-weather fallback evaluation. |
| 4D Radar-Camera Occupancy | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | prototype | 4D Radar-Camera Occupancy is rated for alternative-sensor perception and adverse-weather fallback evaluation. |
| 4DMOS | Learning: ★★★☆☆ Deployment: ★★★★☆ | modern-core | prototype | Extends LiDAR motion segmentation with temporal 4D reasoning. |
| Adverse-Weather Radar-LiDAR 3D Detection | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | prototype | Adverse-Weather Radar-LiDAR 3D Detection is rated for alternative-sensor perception and adverse-weather fallback evaluation. |
| AdverseNet | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | pilot-proven | AdverseNet is rated for cleaning, stress testing, or failure detection in degraded perception conditions. |
| AevaScenes | Learning: ★★★☆☆ Deployment: ★★★★☆ | reference | fielded-pattern | AevaScenes is rated as a benchmark or dataset reference for perception robustness and validation coverage. |
| AIDE | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | pilot-proven | AIDE is rated for operational perception validation, calibration, or safety-screening workflows. |
| Classical LiDAR Outlier Removal | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | pilot-proven | Classical LiDAR Outlier Removal is rated for cleaning, stress testing, or failure detection in degraded perception conditions. |
| CVFusion | Learning: ★★★☆☆ Deployment: ★★★★☆ | modern-core | prototype | Important radar-camera fusion method for degraded visual conditions. |
| DenoiseCP-Net | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | pilot-proven | DenoiseCP-Net is rated for cleaning, stress testing, or failure detection in degraded perception conditions. |
| Ev-3DOD | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | prototype | Ev-3DOD is rated for alternative-sensor perception and adverse-weather fallback evaluation. |
| EvOcc | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | prototype | EvOcc is rated for alternative-sensor perception and adverse-weather fallback evaluation. |
| Fail2Drive | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | pilot-proven | Fail2Drive is rated for operational perception validation, calibration, or safety-screening workflows. |
| K-Radar | Learning: ★★★☆☆ Deployment: ★★★★☆ | modern-core | fielded-pattern | Key 4D radar dataset and benchmark for all-weather perception evaluation. |
| LASP | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | pilot-proven | LASP is rated for operational perception validation, calibration, or safety-screening workflows. |
| LiDAR Weather Artifact Removal | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | pilot-proven | LiDAR Weather Artifact Removal is rated for cleaning, stress testing, or failure detection in degraded perception conditions. |
| LIORNet | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | pilot-proven | LIORNet is rated for cleaning, stress testing, or failure detection in degraded perception conditions. |
| LiSnowNet | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | pilot-proven | LiSnowNet is rated for cleaning, stress testing, or failure detection in degraded perception conditions. |
| M-detector LiDAR Point-Stream MED | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | pilot-proven | M-detector LiDAR Point-Stream MED is rated for operational perception validation, calibration, or safety-screening workflows. |
| MoME | Learning: ★★★☆☆ Deployment: ★★★★☆ | modern-core | prototype | Useful resilient fusion pattern for adverse sensor failure cases. |
| MSC-Bench | Learning: ★★★☆☆ Deployment: ★★★★☆ | reference | fielded-pattern | MSC-Bench is rated as a benchmark or dataset reference for perception robustness and validation coverage. |
| MultiCorrupt | Learning: ★★★☆☆ Deployment: ★★★★☆ | reference | fielded-pattern | MultiCorrupt is rated as a benchmark or dataset reference for perception robustness and validation coverage. |
| Occluded nuScenes | Learning: ★★★☆☆ Deployment: ★★★★☆ | reference | fielded-pattern | Occluded nuScenes is rated as a benchmark or dataset reference for perception robustness and validation coverage. |
| POD FMCW LiDAR Predictive Detection | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | prototype | POD FMCW LiDAR Predictive Detection is rated for alternative-sensor perception and adverse-weather fallback evaluation. |
| ProOOD | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | pilot-proven | ProOOD is rated for cleaning, stress testing, or failure detection in degraded perception conditions. |
| QuantV2X | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | prototype | QuantV2X is rated for deployment-oriented cooperative perception because it quantizes model execution and transmitted V2X feature messages. |
| RaCFormer | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | prototype | RaCFormer is rated for alternative-sensor perception and adverse-weather fallback evaluation. |
| RC-AutoCalib | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | pilot-proven | RC-AutoCalib is rated for operational perception validation, calibration, or safety-screening workflows. |
| RobuRCDet | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | prototype | RobuRCDet is rated for alternative-sensor perception and adverse-weather fallback evaluation. |
| S2R-Bench | Learning: ★★★☆☆ Deployment: ★★★★☆ | reference | fielded-pattern | S2R-Bench is rated as a benchmark or dataset reference for perception robustness and validation coverage. |
| SLiDE | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | pilot-proven | SLiDE is rated for cleaning, stress testing, or failure detection in degraded perception conditions. |
| Sparse4D | Learning: ★★★☆☆ Deployment: ★★★★☆ | modern-core | prototype | Practical sparse-query direction for camera 3D detection and tracking. |
| TripleMixer | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | pilot-proven | TripleMixer is rated for cleaning, stress testing, or failure detection in degraded perception conditions. |
| V2X-Radar | Learning: ★★★☆☆ Deployment: ★★★★☆ | deployment-pattern | prototype | V2X-Radar is rated for alternative-sensor perception and adverse-weather fallback evaluation. |
| DETR4D | Learning: ★★★★☆ Deployment: ★★★☆☆ | modern-core | prototype | DETR4D is rated as a foundational sparse-query camera 3D detection baseline for temporal multi-view perception. |
| FlatFormer | Learning: ★★★★☆ Deployment: ★★★☆☆ | modern-core | prototype | FlatFormer's flattened window attention is the first point-cloud transformer to run real-time on edge GPUs — an efficiency reference for embedded LiDAR backbones. |
| GaussTR | Learning: ★★★★☆ Deployment: ★★★☆☆ | frontier | prototype | GaussTR is rated for foundation-model-aligned, open-vocabulary semantic occupancy research with released code. |
| GS-Occ3D | Learning: ★★★★☆ Deployment: ★★★☆☆ | frontier | prototype | GS-Occ3D is rated for scalable vision-only occupancy reconstruction and Gaussian-surfel label curation. |
| LaserMix | Learning: ★★★★☆ Deployment: ★★★☆☆ | modern-core | prototype | LaserMix is the LiDAR-aware consistency-regularization augmentation that unlocks label-efficient semi-supervised 3D segmentation. |
| OctFormer | Learning: ★★★★☆ Deployment: ★★★☆☆ | modern-core | prototype | OctFormer is an octree-based transformer for efficient semantic segmentation of large 3D point clouds and registered maps. |
| OpenScene | Learning: ★★★★☆ Deployment: ★★★☆☆ | frontier | research | OpenScene is the canonical open-vocabulary 3D segmentation method — text-query the labeled cloud without fixed-taxonomy training. |
| Point-Cloud Mamba / SSM Backbones | Learning: ★★★★☆ Deployment: ★★★☆☆ | frontier | research | Emerging linear-complexity SSM backbones for large point-cloud segmentation and map-scale labeling, with limited production evidence. |
| PointNeXt | Learning: ★★★★☆ Deployment: ★★★☆☆ | classic-baseline | prototype | PointNeXt is the modernized PointNet++ that shows most of the architecture-vs-transformer gap was a training-recipe artifact. |
| PolarMix | Learning: ★★★★☆ Deployment: ★★★☆☆ | modern-core | prototype | PolarMix is the polar-coordinate mixing augmentation that improves single-scan LiDAR segmentation and rebalances rare classes via instance rotate-paste. |
| RangeFormer | Learning: ★★★★☆ Deployment: ★★★☆☆ | modern-core | prototype | RangeFormer is the range-image transformer that first made range-image projection competitive with voxel and fusion methods on LiDAR semantic segmentation. |
| SphereFormer | Learning: ★★★★☆ Deployment: ★★★☆☆ | modern-core | prototype | SphereFormer's radial-window attention targets the sparse far-range density gap in single-scan LiDAR semantic segmentation. |
| VOGS-CP | Learning: ★★★★☆ Deployment: ★★★☆☆ | frontier | prototype | VOGS-CP is rated for collaborative Gaussian semantic occupancy because it exchanges sparse 3D semantic Gaussian primitives instead of dense voxel or planar V2X features. |
| 3D-AVS | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | 3D-AVS is rated for open-world perception, annotation leverage, and long-tail validation workflows. |
| 3D-OutDet | Learning: ★★★☆☆ Deployment: ★★★☆☆ | modern-core | prototype | 3D-OutDet is rated as a supporting perception method for autonomy-stack triage and follow-up reading. |
| Clipomaly | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | Useful anomaly-detection reference for long-tail discovery workflows. |
| CoHFF | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | prototype | CoHFF is rated for cooperative perception and infrastructure-assisted sensing evaluation. |
| CoInfra | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | prototype | CoInfra is rated for cooperative perception and infrastructure-assisted sensing evaluation. |
| CoopTrack | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | prototype | CoopTrack is rated for cooperative perception and infrastructure-assisted sensing evaluation. |
| CoSDH | Learning: ★★★☆☆ Deployment: ★★★☆☆ | modern-core | prototype | CoSDH is rated as a supporting perception method for autonomy-stack triage and follow-up reading. |
| DetAny3D | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | DetAny3D is rated for open-world perception, annotation leverage, and long-tail validation workflows. |
| DistillNeRF | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | DistillNeRF is rated for neural scene representation learning and simulation-oriented perception research. |
| DrivingGaussian | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | DrivingGaussian is rated for neural scene representation learning and simulation-oriented perception research. |
| ForeSight | Learning: ★★★☆☆ Deployment: ★★★☆☆ | modern-core | prototype | ForeSight is rated as a supporting perception method for autonomy-stack triage and follow-up reading. |
| GaussianFormer | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | GaussianFormer is rated for neural scene representation learning and simulation-oriented perception research. |
| HoloVIC | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | prototype | HoloVIC is rated for cooperative perception and infrastructure-assisted sensing evaluation. |
| HUGS Urban Gaussians | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | HUGS Urban Gaussians is rated for neural scene representation learning and simulation-oriented perception research. |
| LOSC | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | LOSC is rated for open-vocabulary LiDAR pseudo-label consolidation, annotation leverage, and map-labeling data-engine workflows. |
| Mosaic3D | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | Mosaic3D is rated for open-vocabulary 3D segmentation, dataset leverage, and long-tail perception validation. |
| OP3Det | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | OP3Det is rated for open-world perception, annotation leverage, and long-tail validation workflows. |
| Open3DTrack | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | Open3DTrack is rated for open-world perception, annotation leverage, and long-tail validation workflows. |
| OpenVox | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | OpenVox is rated for open-world perception, annotation leverage, and long-tail validation workflows. |
| OVAD And OVODA Open-Vocabulary 3D Attributes | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | OVAD And OVODA Open-Vocabulary 3D Attributes is rated for open-world perception, annotation leverage, and long-tail validation workflows. |
| OW-OVD | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | OW-OVD is rated for open-world perception, annotation leverage, and long-tail validation workflows. |
| RCooper | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | fielded-pattern | Cooperative-perception dataset relevant to infrastructure-assisted sensing. |
| S2M | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | S2M is rated for open-world perception, annotation leverage, and long-tail validation workflows. |
| SAM 3 | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | SAM 3 is rated for open-world perception, annotation leverage, and long-tail validation workflows. |
| SAM4D | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | SAM4D is rated for open-world perception, annotation leverage, and long-tail validation workflows. |
| SAMFusion | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | SAMFusion is rated for open-world perception, annotation leverage, and long-tail validation workflows. |
| SOAC | Learning: ★★★☆☆ Deployment: ★★★☆☆ | modern-core | prototype | SOAC is rated as a supporting perception method for autonomy-stack triage and follow-up reading. |
| SpaCeFormer | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | SpaCeFormer is rated for fast proposal-free open-vocabulary 3D instance segmentation and indoor 3D mask-text data generation. |
| SparseCoop | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | prototype | SparseCoop is rated for cooperative sparse-query detection and tracking because it avoids dense BEV feature exchange with kinematic-grounded queries. |
| SplatAD | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | SplatAD is rated for neural scene representation learning and simulation-oriented perception research. |
| SplatFlow | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | research | SplatFlow is rated for neural scene representation learning and simulation-oriented perception research. |
| TacoDepth | Learning: ★★★☆☆ Deployment: ★★★☆☆ | modern-core | prototype | TacoDepth is rated as a supporting perception method for autonomy-stack triage and follow-up reading. |
| V2X-ReaLO | Learning: ★★★☆☆ Deployment: ★★★☆☆ | frontier | prototype | V2X-ReaLO is rated for cooperative perception and infrastructure-assisted sensing evaluation. |
| WildDet3D | Learning: ★★★☆☆ Deployment: ★★★☆☆ | modern-core | prototype | WildDet3D is rated as a supporting perception method for autonomy-stack triage and follow-up reading. |
Domain Fit Guidance
Generic method pages should use Domain Fit, not Airside Fit, as the default deployment lens. Use three to six compact rows or bullets rather than a large matrix.
| Domain | Fit | Note |
|---|---|---|
| Road AV | strong / conditional / weak / insufficient evidence | State whether the method has road-scale evidence, actor coverage, and runtime maturity. |
| Airside | strong / conditional / weak / insufficient evidence | Include apron, GSE, FOD, aircraft-proximity, and weather relevance only when supported by the method evidence. |
| Warehouse / logistics yard / port / mining / construction / agriculture / delivery robot / outdoor campus | strong / conditional / weak / insufficient evidence | Add only the domains where the method assumptions or validation signals materially transfer. |
Airside-specific pages may stay airside-first, but generic pages should not make airside the only deployment lens.
How to Use This Library
For loss and residual foundations behind perception methods, use 3D Object Detection Losses and Assignment for detector training objectives and Robust Losses and M-Estimators for outlier-heavy geometric residuals that connect perception outputs to calibration, tracking, and SLAM.
File Boundary Rules
| Rule | Practical meaning |
|---|---|
| One file, one method | A page should not bundle multiple unrelated methods just because they share a modality. If two papers solve the same exact technique lineage, the page can compare versions, but the title must still name the primary method. |
| Overview pages link out | Existing files such as BEV Encoding Architectures, Streaming Temporal Perception, and Infrastructure Cooperative Perception should summarize families and point here for method-level details. |
| Benchmarks count as methods when they shape evaluation | Pages such as MSC-Bench, S2R-Bench, LASP, OpenAD, and Fail2Drive deserve first-class treatment because they define what a deployment team measures. |
| Domain fit is mandatory for generic pages | Generic method pages should state the domains where the method is a strong, conditional, weak, or insufficient-evidence fit. Airside-specific pages may use a transfer note instead. |
| Sources stay close to claims | Each method page must include primary paper, project, dataset, or repository links so future refreshes can verify claims quickly. |
Standard Page Shape
Each method page should include:
- What the method is.
- Core technical idea.
- Inputs, outputs, and model/data assumptions.
- Architecture or pipeline.
- Training/evaluation setup and benchmark signals.
- Strengths.
- Failure modes and deployment risks.
- Domain Fit, or an airside transfer note when the page is explicitly airside-specific.
- Implementation notes.
- Sources.
Relationship to the Perception Stack
| Existing synthesis page | Method-library role |
|---|---|
| BEV Encoding Architectures | Explains the BEV design space, then links to BEVDet/BEVDepth/BEVStereo/SOLOFusion and camera occupancy methods. |
| Camera-Only Degraded Perception | Uses camera BEV, occupancy, depth, and open-vocabulary method pages to define fallback modes. |
| LiDAR Semantic Segmentation | Summarizes segmentation architecture choices, then links to LiDAR-MOS, 4DMOS, SegNet4D, Mask4D, MotionSeg3D, MambaMOS, neural scene-flow priors, and HeLiMOS-style evaluation. |
| Aggregated-Map Semantic Segmentation | Offline map-scale segmentation pipeline with a compact proxy/input/training selector; anchors the segmentation-backbone and label-consolidation method pages and the architecture reading path below: MinkowskiNet, SPVCNN, KPConv, PointNeXt, RandLA-Net, Cylinder3D, SphereFormer, WaffleIron, SalsaNext, RangeFormer, FRNet, Point Transformer V3, Point-Cloud Mamba / SSM Backbones, FlatFormer, OctFormer, Superpoint Transformer, 2DPASS, OpenScene, Mosaic3D, and LOSC. |
| 3D Segmentation Class Taxonomy Design | Class-taxonomy deep-dive companion to the aggregated-map page: stuff/things, granularity, hierarchical taxonomies, cross-dataset harmonization, open-set handling, and airside map taxonomy. |
| 3D Segmentation Training Paradigms | Training-architecture comparison across supervised, multi-dataset, SSL, distillation, weak, semi-supervised, domain-adaptation, synthetic, PEFT, active-learning, and auto-label patterns. |
| Segmentation Post-Processing and Label Refinement | Post-processing companion for CRF/geometric smoothing, TTA, ensembling, training-free and learned panoptic extraction, calibration, multi-pass fusion, tile stitching, and production-stack assembly. |
| Large-Scale 3D Segmentation Tiling and Throughput | Tiling, stitching, and throughput engineering for map-scale pipelines: partition strategies, halo selection, logit merging, normalization-layer seam artifacts, batched inference, and cost modeling. |
| LiDAR Artifact Removal Techniques | Synthesizes learned denoisers, classical filters, weather artifact handling, ghost/multipath failures, validation, datasets, and map-cleaning links. |
| Streaming Temporal Perception | Connects StreamMOS, 4DSegStreamer, MotionSeg3D, MambaMOS, LASP, sparse-query detection, scene flow, and temporal occupancy into a runtime stack. |
| Open-Vocabulary and Zero-Shot Detection | Stays as the broad open-vocabulary primer; OpenAD, OP3Det, WildDet3D, DetAny3D, OW-OVD, Clipomaly, S2M, SAM 3, and LOSC get individual pages here. |
| Infrastructure Cooperative Perception | Synthesizes V2X deployment tradeoffs; RCooper, HoloVIC, CoInfra, V2X-ReaLO, CoHFF, VOGS-CP, CoSDH, CoopTrack, SparseCoop, QuantV2X, and TruckV2X live here as atomic references. |
| Production Perception Systems | Uses this library as the evidence base for validation matrices, degradation policies, and sensor-suite decisions. |
Aggregated-Map Architecture Reading Path
Use this as the method-library entry for the architecture sections of Aggregated-Map Semantic Segmentation and the training-route details in 3D Segmentation Training Paradigms. A release-map project should not pick a single "best" backbone from a leaderboard. It should select an architecture lane by source-map quality, product mode, modality contract, thin-class risk, tiling cost, and what evidence the team can reproduce on owned non-road district maps.
| Architecture lane | Method pages | Use when | Advantages | Controls and disadvantages |
|---|---|---|---|---|
| Sparse-voxel production anchor | MinkowskiNet, SPVCNN | The first runtime semantic-map or training-export product needs mature tooling, predictable batching, and a LiDAR-only or colorized-cloud input contract. | Strong accuracy/throughput balance, common production baseline, clean voxel/tile packaging, easier acceleration with sparse-conv backends. | Voxel size trades thin-structure recall against memory; audit curb, marking, cable, pole, fence, and facade-detail loss, plus release-state seam confusion. |
| Point-conv geometry-faithful baseline | KPConv, RandLA-Net, PointNeXt | The bake-off needs a raw-geometry reference for MLS/ALS survey clouds, thin classes, facade details, and utility or non-road infrastructure. | Preserves point geometry without voxelization loss; long survey-industry lineage; useful for sanity-checking sparse-conv detail loss. | Neighbor search and sphere sampling are expensive; RandLA-style random sampling can drop rare classes unless the seed policy is rare-class aware. |
| Superpoint graph map-scale lane | Superpoint Transformer | Airport, campus, port, utility-corridor, terminal-frontage, or facade maps need larger context and fewer hard tile seams. | Superpoint partitioning reuses geometric structure, reduces memory, exposes whole-scene context, and fits scarce-label regimes well. | Partition quality caps accuracy; boundary recall, under-segmentation, over-segmentation, and partition-version evidence need their own sidecar. |
| Serialized transformer / foundation-model ceiling | Point Transformer V3, OctFormer, SphereFormer | The project has public or internal pre-training, enough GPU budget, and wants an accuracy-ceiling comparison after the conservative baselines are stable. | Highest ceiling under good pre-training and tiling; serialized attention and octree/window variants carry more context than plain local convolution. | Heavier training and inference; warmup, optimizer stability, tile halo, no-clipping boundary policy, and ensemble cost must be held constant before claiming a real gain. |
| Projection and standard-op deployment lane | WaffleIron, SalsaNext, RangeFormer, FRNet, Cylinder3D, FlatFormer | The team needs a dependency-light baseline, an embedded-friendly single-scan sibling, or a fast projection model for color/intensity planes. | Standard dense ops are easier to implement, optimize, and deploy; good for fast ablations and single-scan/map taxonomy alignment. | Projection resolution discards some 3D detail; not a default release-map backbone unless thin-object, vertical-structure, and occlusion audits pass. |
| Image-distilled and open-vocabulary label lane | 2DPASS, OpenScene, Mosaic3D, LOSC | Calibrated imagery, 2D foundation models, or open-vocabulary labels can improve supervised coverage, reviewer triage, or long-tail class discovery. | Adds image supervision without necessarily adding an image dependency at release time; helps candidate labels for rare assets and semantic gaps. | Candidate and pseudo labels are not release truth; projection QA, reviewer promotion, taxonomy mapping, and manifest permissions decide whether labels can be consumed. |
| SSM / Mamba efficiency frontier | Point-Cloud Mamba / SSM Backbones | A P3/P4 experiment needs larger context per tile or lower sequence cost after mature baselines define the target. | Linear token cost is structurally attractive for registered maps where context radius and tile count dominate cost. | Research-stage for release maps; ordering, rotation, density shift, local geometry recovery, seam stability, and non-road transfer need explicit audits before production use. |
Minimum bake-off contract: compare the lanes above only under the same source-map acceptance package, taxonomy, release-state labels, tile manifest, halo policy, train/validation/test split, optimizer budget, augmentation schedule, and modality lane. Report semantic mIoU, class-balanced mIoU, rare/thin-class recall, boundary F1, seam disagreement, calibration error, throughput, cost per km2 or per 100M points, release-state confusion, false-permanent rate, and false-deletion rate. This prevents the common failure where a higher leaderboard mIoU hides an unusable runtime dependency, contaminated training export, or brittle non-road district transfer.
Expansion Backlog
The first waves focused on methods already identified as P0/P1 in the Perception Coverage Audit. The 2026-05-09 loops promoted SplatAD, GaussianFormer, GaussianOcc, streaming Gaussian occupancy, Cam4DOcc, StreamingFlow, Sparse4D, TacoDepth, RaCFormer, LIORNet, learned LiDAR desnowing/denoising, broad artifact removal, classical outlier filtering, MotionSeg3D, MambaMOS, neural scene-flow priors, CVFusion, 4D radar-camera occupancy, POD/FMCW LiDAR, DrivingGaussian, HUGS, SplatFlow, DistillNeRF, TrackOcc, cross-domain scene flow, LiDAR-camera occupancy fusion, dynamic occupancy/free-space, radar-LiDAR adverse-weather detection, RobuRCDet, SAMFusion, spatiotemporal memory occupancy flow, OVAD/OVODA, and open-vocabulary panoptic occupancy into atomic files. Future waves should split remaining grouped rows into atomic pages, especially:
- VEON, ProOOD, and SA-Occ. EvOcc is already promoted; DR-REMOVER and ExelMap are routed through DR-REMOVER and ExelMap rather than duplicate perception pages.
- Drive-OccWorld and DFIT-OccWorld where they need separate world-model or planning-facing treatment beyond the dynamic occupancy page.
- DySS is routed through the sparse-query and streaming-temporal overviews pending official code or broader adoption; remaining sparse-query or end-to-end follow-ons should only become atomic pages when they add a non-duplicative method boundary beyond SparseBEV, DETR4D, Sparse4D, ForeSight, SparseDrive, or DiffusionDrive.
- LinkOcc and related 2025-2026 occupancy follow-ons; GaussianFlowOcc, GaussTR, and GS-Occ3D are now first-class pages, while missing-view occupancy routes through M2-Occ, temporal radar-camera occupancy routes through TEOcc, and Gaussian-rendered occupancy routes through GaussRender unless a distinct method needs its own page.
- CoDS, JigsawComm, and collaborative Gaussian follow-ons such as GSCOOP. VOGS-CP is already promoted as the collaborative Gaussian semantic-occupancy reference, SparseCoop is already promoted as the sparse cooperative-query reference, QuantV2X is already promoted as the cooperative quantization reference, and TruckV2X is promoted as the truck-centered cooperative dataset reference.
- Residual FOD OOD-vs-detector benchmark planning and airside-specific dust/de-icing-mist/steam/glycol/wet-apron datasets; SpaCeFormer now covers fast proposal-free open-vocabulary 3D instance segmentation, EmbodiedScan/MMScan covers embodied robotics 3D perception benchmarks, and DriveBench stays discoverable through the VLA/VLM reliability benchmark page rather than duplicating it here.
Sources
- Perception coverage audit and backlog: coverage-audit-2026.md
- Existing perception synthesis index: Research Index - Perception