Skip to content

Perception Method Library Overview

This directory is the method-level perception library. Each page should represent one technique, method, benchmark, or dataset-backed evaluation primitive. Broad synthesis pages in 30-autonomy-stack/perception/overview/ remain useful for system design, but this library is where individual methods get enough space for architecture, data, benchmarks, failure modes, deployment fit, Domain Fit, transfer notes for explicitly scoped ODDs, and sources.

Priority Ratings

Priority ratings are editorial reading and deployment triage signals. Learning answers what to read early for general autonomy understanding. Deployment answers what to evaluate early for AV deployment in the tagged context; it is not a certification, product-readiness, or all-domain average claim. If a method's deployment score is driven by a specific domain or stack role, the reason text should name that context.

MethodRatingStageMaturityReason
MinkowskiNetLearning: ★★★★★
Deployment: ★★★★★
classic-baselinefielded-patternMinkowskiNet sparse-convolution U-Nets are the deployed industry-baseline backbone for 3D semantic segmentation, including offline aggregated-map labeling.
Availability-Aware Sensor FusionLearning: ★★★★☆
Deployment: ★★★★★
deployment-patternprototypeDirectly targets sensor degradation and availability-aware fusion.
LiDAR-MOSLearning: ★★★★☆
Deployment: ★★★★★
deployment-patternprototypeMoving-object segmentation is central to map hygiene and dynamic-scene handling.
KPConvLearning: ★★★★★
Deployment: ★★★★☆
classic-baselinefielded-patternKPConv is the point-based-convolution workhorse for large-scale MLS/ALS point cloud semantic segmentation and a proven baseline for aggregated-map labeling.
Point Transformer V3Learning: ★★★★★
Deployment: ★★★★☆
modern-coreprototypePoint Transformer V3 is the leading serialized-attention 3D backbone for large-scale point cloud semantic segmentation, including offline aggregated-map labeling.
RandLA-NetLearning: ★★★★★
Deployment: ★★★★☆
classic-baselinefielded-patternRandLA-Net is the efficient large-scale point-cloud segmentation baseline, designed for million-point clouds and used across MLS survey workflows.
2DPASSLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototype2DPASS is the reference training scheme for image-assisted LiDAR semantic segmentation with LiDAR-only inference.
4DSegStreamerLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototype4DSegStreamer is rated for motion segmentation, scene flow, or dynamic-object perception workflows.
AutoOccLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeAutoOcc is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks.
BEVDepthLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeImportant depth-aware BEV bridge for camera-only 3D perception.
BEVDetLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeBaseline camera BEV detector that organizes many later BEV methods.
BEVStereoLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeBEVStereo is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks.
Cam4DOccLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeCam4DOcc is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks.
Conformal BoxesLearning: ★★★★☆
Deployment: ★★★★☆
deployment-patternprototypePractical uncertainty wrapper for detection risk and release gates.
Cross-Domain LiDAR Scene FlowLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeCross-Domain LiDAR Scene Flow is rated for motion segmentation, scene flow, or dynamic-object perception workflows.
Cylinder3DLearning: ★★★★☆
Deployment: ★★★★☆
classic-baselineprototypeCylinder3D is the cylindrical sparse-voxel baseline for single-scan spinning-LiDAR semantic segmentation.
DepthOccLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeDepthOcc is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks.
DiffusionDriveLearning: ★★★★☆
Deployment: ★★★★☆
frontierprototypeTruncated diffusion policy that makes diffusion-based end-to-end driving fast enough for real-time use.
Dynamic Occupancy FreespaceLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeDynamic Occupancy Freespace is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks.
FlashOccLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeFlashOcc is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks.
FRNetLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeFRNet is the frustum-range network that matches RangeFormer's accuracy at roughly 5x the speed — the current real-time range-image SOTA.
GaussianFlowOccLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeGaussianFlowOcc is rated for sparse, weakly supervised semantic occupancy with explicit temporal Gaussian flow.
GaussianOccLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeGaussianOcc is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks.
GaussRenderLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeGaussRender is rated for improving 3D occupancy geometry through efficient projective Gaussian-rendering supervision.
GraphBEVLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeGraphBEV is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks.
InsMOSLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeInsMOS is rated for motion segmentation, scene flow, or dynamic-object perception workflows.
Instantaneous Motion PerceptionLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeInstantaneous Motion Perception is rated for motion segmentation, scene flow, or dynamic-object perception workflows.
LiDAR-Camera Occupancy FusionLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeLiDAR-Camera Occupancy Fusion is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks.
M2-OccLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeM2-Occ is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks.
MambaMOSLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeMambaMOS is rated for motion segmentation, scene flow, or dynamic-object perception workflows.
Mask4DLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeMask4D is rated for motion segmentation, scene flow, or dynamic-object perception workflows.
MotionSeg3DLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeMotionSeg3D is rated for motion segmentation, scene flow, or dynamic-object perception workflows.
Neural Scene Flow PriorsLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeNeural Scene Flow Priors is rated for motion segmentation, scene flow, or dynamic-object perception workflows.
Open-Vocabulary Panoptic OccupancyLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeOpen-Vocabulary Panoptic Occupancy is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks.
OpenADLearning: ★★★★☆
Deployment: ★★★★☆
modern-corefielded-patternOpen-world benchmark for corner cases and unseen categories.
RadarPillarsLearning: ★★★★☆
Deployment: ★★★★☆
classic-baselineprototypeCore radar-native detection baseline for weather-robust perception.
RenderOccLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeRenderOcc is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks.
SalsaNextLearning: ★★★★☆
Deployment: ★★★★☆
classic-baselinepilot-provenSalsaNext is the canonical real-time range-image LiDAR semantic segmenter — the fast, embedded-friendly projection baseline.
SegNet4DLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeSegNet4D is rated for motion segmentation, scene flow, or dynamic-object perception workflows.
SelfOccLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeSelfOcc is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks.
SOLOFusionLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeSOLOFusion is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks.
SparseBEVLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeSparseBEV is rated for efficient sparse-query multi-camera 3D detection where dense BEV memory is costly.
SparseDriveLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeFully sparse end-to-end driving stack that unifies detection, tracking, online mapping, and motion planning.
SparseOccLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeSparseOcc is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks.
Spatiotemporal Memory Occupancy FlowLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeSpatiotemporal Memory Occupancy Flow is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks.
SPVCNNLearning: ★★★★☆
Deployment: ★★★★☆
classic-baselineprototypeSPVCNN's sparse point-voxel convolution is the efficient point-voxel-hybrid baseline for real-time LiDAR semantic segmentation.
Streaming Gaussian OccupancyLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeStreaming Gaussian Occupancy is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks.
StreamingFlowLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeStreamingFlow is rated for motion segmentation, scene flow, or dynamic-object perception workflows.
StreamMOSLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeStreamMOS is rated for motion segmentation, scene flow, or dynamic-object perception workflows.
Superpoint TransformerLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeSuperpoint Transformer is rated for large-scale 3D semantic segmentation of registered point clouds where whole-scene context and very small models matter.
SurroundOccLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeFoundational camera occupancy reference for planning-facing perception.
TEOccLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeTEOcc is rated for radar-camera semantic occupancy and temporal robustness under degraded perception conditions.
TPVFormerLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeTPVFormer is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks.
TrackOccLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeTrackOcc is rated for BEV, occupancy, or freespace modeling that feeds planning-facing autonomy stacks.
WaffleIronLearning: ★★★★☆
Deployment: ★★★★☆
modern-coreprototypeWaffleIron is a deliberately simple LiDAR segmentation backbone built from standard dense 2D convolutions, easy to implement and deploy.
3D-KNN Blind-Spot DesnowingLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternpilot-proven3D-KNN Blind-Spot Desnowing is rated for cleaning, stress testing, or failure detection in degraded perception conditions.
4D Radar Road Boundaries and FreespaceLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternprototype4D Radar Road Boundaries and Freespace is rated for alternative-sensor perception and adverse-weather fallback evaluation.
4D Radar-Camera OccupancyLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternprototype4D Radar-Camera Occupancy is rated for alternative-sensor perception and adverse-weather fallback evaluation.
4DMOSLearning: ★★★☆☆
Deployment: ★★★★☆
modern-coreprototypeExtends LiDAR motion segmentation with temporal 4D reasoning.
Adverse-Weather Radar-LiDAR 3D DetectionLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternprototypeAdverse-Weather Radar-LiDAR 3D Detection is rated for alternative-sensor perception and adverse-weather fallback evaluation.
AdverseNetLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternpilot-provenAdverseNet is rated for cleaning, stress testing, or failure detection in degraded perception conditions.
AevaScenesLearning: ★★★☆☆
Deployment: ★★★★☆
referencefielded-patternAevaScenes is rated as a benchmark or dataset reference for perception robustness and validation coverage.
AIDELearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternpilot-provenAIDE is rated for operational perception validation, calibration, or safety-screening workflows.
Classical LiDAR Outlier RemovalLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternpilot-provenClassical LiDAR Outlier Removal is rated for cleaning, stress testing, or failure detection in degraded perception conditions.
CVFusionLearning: ★★★☆☆
Deployment: ★★★★☆
modern-coreprototypeImportant radar-camera fusion method for degraded visual conditions.
DenoiseCP-NetLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternpilot-provenDenoiseCP-Net is rated for cleaning, stress testing, or failure detection in degraded perception conditions.
Ev-3DODLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternprototypeEv-3DOD is rated for alternative-sensor perception and adverse-weather fallback evaluation.
EvOccLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternprototypeEvOcc is rated for alternative-sensor perception and adverse-weather fallback evaluation.
Fail2DriveLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternpilot-provenFail2Drive is rated for operational perception validation, calibration, or safety-screening workflows.
K-RadarLearning: ★★★☆☆
Deployment: ★★★★☆
modern-corefielded-patternKey 4D radar dataset and benchmark for all-weather perception evaluation.
LASPLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternpilot-provenLASP is rated for operational perception validation, calibration, or safety-screening workflows.
LiDAR Weather Artifact RemovalLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternpilot-provenLiDAR Weather Artifact Removal is rated for cleaning, stress testing, or failure detection in degraded perception conditions.
LIORNetLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternpilot-provenLIORNet is rated for cleaning, stress testing, or failure detection in degraded perception conditions.
LiSnowNetLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternpilot-provenLiSnowNet is rated for cleaning, stress testing, or failure detection in degraded perception conditions.
M-detector LiDAR Point-Stream MEDLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternpilot-provenM-detector LiDAR Point-Stream MED is rated for operational perception validation, calibration, or safety-screening workflows.
MoMELearning: ★★★☆☆
Deployment: ★★★★☆
modern-coreprototypeUseful resilient fusion pattern for adverse sensor failure cases.
MSC-BenchLearning: ★★★☆☆
Deployment: ★★★★☆
referencefielded-patternMSC-Bench is rated as a benchmark or dataset reference for perception robustness and validation coverage.
MultiCorruptLearning: ★★★☆☆
Deployment: ★★★★☆
referencefielded-patternMultiCorrupt is rated as a benchmark or dataset reference for perception robustness and validation coverage.
Occluded nuScenesLearning: ★★★☆☆
Deployment: ★★★★☆
referencefielded-patternOccluded nuScenes is rated as a benchmark or dataset reference for perception robustness and validation coverage.
POD FMCW LiDAR Predictive DetectionLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternprototypePOD FMCW LiDAR Predictive Detection is rated for alternative-sensor perception and adverse-weather fallback evaluation.
ProOODLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternpilot-provenProOOD is rated for cleaning, stress testing, or failure detection in degraded perception conditions.
QuantV2XLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternprototypeQuantV2X is rated for deployment-oriented cooperative perception because it quantizes model execution and transmitted V2X feature messages.
RaCFormerLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternprototypeRaCFormer is rated for alternative-sensor perception and adverse-weather fallback evaluation.
RC-AutoCalibLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternpilot-provenRC-AutoCalib is rated for operational perception validation, calibration, or safety-screening workflows.
RobuRCDetLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternprototypeRobuRCDet is rated for alternative-sensor perception and adverse-weather fallback evaluation.
S2R-BenchLearning: ★★★☆☆
Deployment: ★★★★☆
referencefielded-patternS2R-Bench is rated as a benchmark or dataset reference for perception robustness and validation coverage.
SLiDELearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternpilot-provenSLiDE is rated for cleaning, stress testing, or failure detection in degraded perception conditions.
Sparse4DLearning: ★★★☆☆
Deployment: ★★★★☆
modern-coreprototypePractical sparse-query direction for camera 3D detection and tracking.
TripleMixerLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternpilot-provenTripleMixer is rated for cleaning, stress testing, or failure detection in degraded perception conditions.
V2X-RadarLearning: ★★★☆☆
Deployment: ★★★★☆
deployment-patternprototypeV2X-Radar is rated for alternative-sensor perception and adverse-weather fallback evaluation.
DETR4DLearning: ★★★★☆
Deployment: ★★★☆☆
modern-coreprototypeDETR4D is rated as a foundational sparse-query camera 3D detection baseline for temporal multi-view perception.
FlatFormerLearning: ★★★★☆
Deployment: ★★★☆☆
modern-coreprototypeFlatFormer's flattened window attention is the first point-cloud transformer to run real-time on edge GPUs — an efficiency reference for embedded LiDAR backbones.
GaussTRLearning: ★★★★☆
Deployment: ★★★☆☆
frontierprototypeGaussTR is rated for foundation-model-aligned, open-vocabulary semantic occupancy research with released code.
GS-Occ3DLearning: ★★★★☆
Deployment: ★★★☆☆
frontierprototypeGS-Occ3D is rated for scalable vision-only occupancy reconstruction and Gaussian-surfel label curation.
LaserMixLearning: ★★★★☆
Deployment: ★★★☆☆
modern-coreprototypeLaserMix is the LiDAR-aware consistency-regularization augmentation that unlocks label-efficient semi-supervised 3D segmentation.
OctFormerLearning: ★★★★☆
Deployment: ★★★☆☆
modern-coreprototypeOctFormer is an octree-based transformer for efficient semantic segmentation of large 3D point clouds and registered maps.
OpenSceneLearning: ★★★★☆
Deployment: ★★★☆☆
frontierresearchOpenScene is the canonical open-vocabulary 3D segmentation method — text-query the labeled cloud without fixed-taxonomy training.
Point-Cloud Mamba / SSM BackbonesLearning: ★★★★☆
Deployment: ★★★☆☆
frontierresearchEmerging linear-complexity SSM backbones for large point-cloud segmentation and map-scale labeling, with limited production evidence.
PointNeXtLearning: ★★★★☆
Deployment: ★★★☆☆
classic-baselineprototypePointNeXt is the modernized PointNet++ that shows most of the architecture-vs-transformer gap was a training-recipe artifact.
PolarMixLearning: ★★★★☆
Deployment: ★★★☆☆
modern-coreprototypePolarMix is the polar-coordinate mixing augmentation that improves single-scan LiDAR segmentation and rebalances rare classes via instance rotate-paste.
RangeFormerLearning: ★★★★☆
Deployment: ★★★☆☆
modern-coreprototypeRangeFormer is the range-image transformer that first made range-image projection competitive with voxel and fusion methods on LiDAR semantic segmentation.
SphereFormerLearning: ★★★★☆
Deployment: ★★★☆☆
modern-coreprototypeSphereFormer's radial-window attention targets the sparse far-range density gap in single-scan LiDAR semantic segmentation.
VOGS-CPLearning: ★★★★☆
Deployment: ★★★☆☆
frontierprototypeVOGS-CP is rated for collaborative Gaussian semantic occupancy because it exchanges sparse 3D semantic Gaussian primitives instead of dense voxel or planar V2X features.
3D-AVSLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearch3D-AVS is rated for open-world perception, annotation leverage, and long-tail validation workflows.
3D-OutDetLearning: ★★★☆☆
Deployment: ★★★☆☆
modern-coreprototype3D-OutDet is rated as a supporting perception method for autonomy-stack triage and follow-up reading.
ClipomalyLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchUseful anomaly-detection reference for long-tail discovery workflows.
CoHFFLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierprototypeCoHFF is rated for cooperative perception and infrastructure-assisted sensing evaluation.
CoInfraLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierprototypeCoInfra is rated for cooperative perception and infrastructure-assisted sensing evaluation.
CoopTrackLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierprototypeCoopTrack is rated for cooperative perception and infrastructure-assisted sensing evaluation.
CoSDHLearning: ★★★☆☆
Deployment: ★★★☆☆
modern-coreprototypeCoSDH is rated as a supporting perception method for autonomy-stack triage and follow-up reading.
DetAny3DLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchDetAny3D is rated for open-world perception, annotation leverage, and long-tail validation workflows.
DistillNeRFLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchDistillNeRF is rated for neural scene representation learning and simulation-oriented perception research.
DrivingGaussianLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchDrivingGaussian is rated for neural scene representation learning and simulation-oriented perception research.
ForeSightLearning: ★★★☆☆
Deployment: ★★★☆☆
modern-coreprototypeForeSight is rated as a supporting perception method for autonomy-stack triage and follow-up reading.
GaussianFormerLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchGaussianFormer is rated for neural scene representation learning and simulation-oriented perception research.
HoloVICLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierprototypeHoloVIC is rated for cooperative perception and infrastructure-assisted sensing evaluation.
HUGS Urban GaussiansLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchHUGS Urban Gaussians is rated for neural scene representation learning and simulation-oriented perception research.
LOSCLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchLOSC is rated for open-vocabulary LiDAR pseudo-label consolidation, annotation leverage, and map-labeling data-engine workflows.
Mosaic3DLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchMosaic3D is rated for open-vocabulary 3D segmentation, dataset leverage, and long-tail perception validation.
OP3DetLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchOP3Det is rated for open-world perception, annotation leverage, and long-tail validation workflows.
Open3DTrackLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchOpen3DTrack is rated for open-world perception, annotation leverage, and long-tail validation workflows.
OpenVoxLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchOpenVox is rated for open-world perception, annotation leverage, and long-tail validation workflows.
OVAD And OVODA Open-Vocabulary 3D AttributesLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchOVAD And OVODA Open-Vocabulary 3D Attributes is rated for open-world perception, annotation leverage, and long-tail validation workflows.
OW-OVDLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchOW-OVD is rated for open-world perception, annotation leverage, and long-tail validation workflows.
RCooperLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierfielded-patternCooperative-perception dataset relevant to infrastructure-assisted sensing.
S2MLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchS2M is rated for open-world perception, annotation leverage, and long-tail validation workflows.
SAM 3Learning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchSAM 3 is rated for open-world perception, annotation leverage, and long-tail validation workflows.
SAM4DLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchSAM4D is rated for open-world perception, annotation leverage, and long-tail validation workflows.
SAMFusionLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchSAMFusion is rated for open-world perception, annotation leverage, and long-tail validation workflows.
SOACLearning: ★★★☆☆
Deployment: ★★★☆☆
modern-coreprototypeSOAC is rated as a supporting perception method for autonomy-stack triage and follow-up reading.
SpaCeFormerLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchSpaCeFormer is rated for fast proposal-free open-vocabulary 3D instance segmentation and indoor 3D mask-text data generation.
SparseCoopLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierprototypeSparseCoop is rated for cooperative sparse-query detection and tracking because it avoids dense BEV feature exchange with kinematic-grounded queries.
SplatADLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchSplatAD is rated for neural scene representation learning and simulation-oriented perception research.
SplatFlowLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierresearchSplatFlow is rated for neural scene representation learning and simulation-oriented perception research.
TacoDepthLearning: ★★★☆☆
Deployment: ★★★☆☆
modern-coreprototypeTacoDepth is rated as a supporting perception method for autonomy-stack triage and follow-up reading.
V2X-ReaLOLearning: ★★★☆☆
Deployment: ★★★☆☆
frontierprototypeV2X-ReaLO is rated for cooperative perception and infrastructure-assisted sensing evaluation.
WildDet3DLearning: ★★★☆☆
Deployment: ★★★☆☆
modern-coreprototypeWildDet3D is rated as a supporting perception method for autonomy-stack triage and follow-up reading.

Domain Fit Guidance

Generic method pages should use Domain Fit, not Airside Fit, as the default deployment lens. Use three to six compact rows or bullets rather than a large matrix.

DomainFitNote
Road AVstrong / conditional / weak / insufficient evidenceState whether the method has road-scale evidence, actor coverage, and runtime maturity.
Airsidestrong / conditional / weak / insufficient evidenceInclude apron, GSE, FOD, aircraft-proximity, and weather relevance only when supported by the method evidence.
Warehouse / logistics yard / port / mining / construction / agriculture / delivery robot / outdoor campusstrong / conditional / weak / insufficient evidenceAdd only the domains where the method assumptions or validation signals materially transfer.

Airside-specific pages may stay airside-first, but generic pages should not make airside the only deployment lens.

How to Use This Library

For loss and residual foundations behind perception methods, use 3D Object Detection Losses and Assignment for detector training objectives and Robust Losses and M-Estimators for outlier-heavy geometric residuals that connect perception outputs to calibration, tracking, and SLAM.

NeedStart here
Camera BEV, occupancy, and freespaceBEVDet, BEVDepth, BEVStereo, SOLOFusion, Sparse4D, SparseBEV, DETR4D, TPVFormer, SurroundOcc, SparseOcc, FlashOcc, DepthOcc, SelfOcc, RenderOcc, GaussRender, GaussTR, GS-Occ3D, LiDAR-Camera Occupancy Fusion, Dynamic Occupancy and Freespace, Spatiotemporal Memory Occupancy Flow
Gaussian, 3DGS, 4DGS, and 4D occupancySplatAD, GaussianFormer, GaussianOcc, GaussianFlowOcc, Streaming Gaussian Occupancy, GS-Occ3D, GaussTR, VOGS-CP, Cam4DOcc, StreamingFlow, DrivingGaussian, HUGS, SplatFlow, DistillNeRF
LiDAR motion, scene flow, and temporal segmentationLiDAR-MOS, 4DMOS, InsMOS, StreamMOS, 4DSegStreamer, SegNet4D, Mask4D, Instantaneous Motion Perception, MotionSeg3D, MambaMOS, Neural Scene Flow Priors, Cross-Domain LiDAR Scene Flow, TrackOcc
3D point cloud semantic segmentation backbonesMinkowskiNet, SPVCNN, KPConv, PointNeXt, RandLA-Net, Cylinder3D, SphereFormer, WaffleIron, SalsaNext, RangeFormer, FRNet, Point Transformer V3, Point-Cloud Mamba / SSM Backbones, FlatFormer, OctFormer, Superpoint Transformer, 2DPASS, OpenScene, Mosaic3D
LiDAR denoising, removal, and adverse weatherLIORNet, LiSnowNet, SLiDE, TripleMixer, 3D-KNN Blind-Spot Desnowing, 3D-OutDet, AdverseNet, DenoiseCP-Net, Classical LiDAR Outlier Removal, LiDAR Weather Artifact Removal
Radar, 4D radar, event, and FMCW perceptionRadarPillars, K-Radar, V2X-Radar, TacoDepth, RaCFormer, CVFusion, 4D Radar-Camera Occupancy, TEOcc, 4D Radar Road Boundaries and Freespace, Adverse-Weather Radar-LiDAR 3D Detection, RobuRCDet, SAMFusion, POD FMCW LiDAR Predictive Detection, Ev-3DOD, AevaScenes
Open-world and open-vocabulary perceptionOpenAD, OP3Det, WildDet3D, DetAny3D, OW-OVD, Clipomaly, S2M, SAM 3, 3D-AVS, Mosaic3D, LOSC, SpaCeFormer, OpenVox, OVAD/OVODA Open-Vocabulary 3D Attributes, Open-Vocabulary Panoptic Occupancy
Robust fusion and perception validationMoME, GraphBEV, SOAC, RC-AutoCalib, ASF, MSC-Bench, MultiCorrupt, S2R-Bench, Occluded nuScenes, Conformal Boxes
Cooperative, online, and data-engine methodsRCooper, HoloVIC, CoInfra, V2X-ReaLO, CoHFF, VOGS-CP, CoSDH, CoopTrack, SparseCoop, QuantV2X, LASP, Fail2Drive, AIDE
Sparse-query and end-to-end drivingSparseBEV, Sparse4D, DETR4D, ForeSight, SparseDrive, DiffusionDrive, SAM4D, Open3DTrack

File Boundary Rules

RulePractical meaning
One file, one methodA page should not bundle multiple unrelated methods just because they share a modality. If two papers solve the same exact technique lineage, the page can compare versions, but the title must still name the primary method.
Overview pages link outExisting files such as BEV Encoding Architectures, Streaming Temporal Perception, and Infrastructure Cooperative Perception should summarize families and point here for method-level details.
Benchmarks count as methods when they shape evaluationPages such as MSC-Bench, S2R-Bench, LASP, OpenAD, and Fail2Drive deserve first-class treatment because they define what a deployment team measures.
Domain fit is mandatory for generic pagesGeneric method pages should state the domains where the method is a strong, conditional, weak, or insufficient-evidence fit. Airside-specific pages may use a transfer note instead.
Sources stay close to claimsEach method page must include primary paper, project, dataset, or repository links so future refreshes can verify claims quickly.

Standard Page Shape

Each method page should include:

  1. What the method is.
  2. Core technical idea.
  3. Inputs, outputs, and model/data assumptions.
  4. Architecture or pipeline.
  5. Training/evaluation setup and benchmark signals.
  6. Strengths.
  7. Failure modes and deployment risks.
  8. Domain Fit, or an airside transfer note when the page is explicitly airside-specific.
  9. Implementation notes.
  10. Sources.

Relationship to the Perception Stack

Existing synthesis pageMethod-library role
BEV Encoding ArchitecturesExplains the BEV design space, then links to BEVDet/BEVDepth/BEVStereo/SOLOFusion and camera occupancy methods.
Camera-Only Degraded PerceptionUses camera BEV, occupancy, depth, and open-vocabulary method pages to define fallback modes.
LiDAR Semantic SegmentationSummarizes segmentation architecture choices, then links to LiDAR-MOS, 4DMOS, SegNet4D, Mask4D, MotionSeg3D, MambaMOS, neural scene-flow priors, and HeLiMOS-style evaluation.
Aggregated-Map Semantic SegmentationOffline map-scale segmentation pipeline with a compact proxy/input/training selector; anchors the segmentation-backbone and label-consolidation method pages and the architecture reading path below: MinkowskiNet, SPVCNN, KPConv, PointNeXt, RandLA-Net, Cylinder3D, SphereFormer, WaffleIron, SalsaNext, RangeFormer, FRNet, Point Transformer V3, Point-Cloud Mamba / SSM Backbones, FlatFormer, OctFormer, Superpoint Transformer, 2DPASS, OpenScene, Mosaic3D, and LOSC.
3D Segmentation Class Taxonomy DesignClass-taxonomy deep-dive companion to the aggregated-map page: stuff/things, granularity, hierarchical taxonomies, cross-dataset harmonization, open-set handling, and airside map taxonomy.
3D Segmentation Training ParadigmsTraining-architecture comparison across supervised, multi-dataset, SSL, distillation, weak, semi-supervised, domain-adaptation, synthetic, PEFT, active-learning, and auto-label patterns.
Segmentation Post-Processing and Label RefinementPost-processing companion for CRF/geometric smoothing, TTA, ensembling, training-free and learned panoptic extraction, calibration, multi-pass fusion, tile stitching, and production-stack assembly.
Large-Scale 3D Segmentation Tiling and ThroughputTiling, stitching, and throughput engineering for map-scale pipelines: partition strategies, halo selection, logit merging, normalization-layer seam artifacts, batched inference, and cost modeling.
LiDAR Artifact Removal TechniquesSynthesizes learned denoisers, classical filters, weather artifact handling, ghost/multipath failures, validation, datasets, and map-cleaning links.
Streaming Temporal PerceptionConnects StreamMOS, 4DSegStreamer, MotionSeg3D, MambaMOS, LASP, sparse-query detection, scene flow, and temporal occupancy into a runtime stack.
Open-Vocabulary and Zero-Shot DetectionStays as the broad open-vocabulary primer; OpenAD, OP3Det, WildDet3D, DetAny3D, OW-OVD, Clipomaly, S2M, SAM 3, and LOSC get individual pages here.
Infrastructure Cooperative PerceptionSynthesizes V2X deployment tradeoffs; RCooper, HoloVIC, CoInfra, V2X-ReaLO, CoHFF, VOGS-CP, CoSDH, CoopTrack, SparseCoop, QuantV2X, and TruckV2X live here as atomic references.
Production Perception SystemsUses this library as the evidence base for validation matrices, degradation policies, and sensor-suite decisions.

Aggregated-Map Architecture Reading Path

Use this as the method-library entry for the architecture sections of Aggregated-Map Semantic Segmentation and the training-route details in 3D Segmentation Training Paradigms. A release-map project should not pick a single "best" backbone from a leaderboard. It should select an architecture lane by source-map quality, product mode, modality contract, thin-class risk, tiling cost, and what evidence the team can reproduce on owned non-road district maps.

Architecture laneMethod pagesUse whenAdvantagesControls and disadvantages
Sparse-voxel production anchorMinkowskiNet, SPVCNNThe first runtime semantic-map or training-export product needs mature tooling, predictable batching, and a LiDAR-only or colorized-cloud input contract.Strong accuracy/throughput balance, common production baseline, clean voxel/tile packaging, easier acceleration with sparse-conv backends.Voxel size trades thin-structure recall against memory; audit curb, marking, cable, pole, fence, and facade-detail loss, plus release-state seam confusion.
Point-conv geometry-faithful baselineKPConv, RandLA-Net, PointNeXtThe bake-off needs a raw-geometry reference for MLS/ALS survey clouds, thin classes, facade details, and utility or non-road infrastructure.Preserves point geometry without voxelization loss; long survey-industry lineage; useful for sanity-checking sparse-conv detail loss.Neighbor search and sphere sampling are expensive; RandLA-style random sampling can drop rare classes unless the seed policy is rare-class aware.
Superpoint graph map-scale laneSuperpoint TransformerAirport, campus, port, utility-corridor, terminal-frontage, or facade maps need larger context and fewer hard tile seams.Superpoint partitioning reuses geometric structure, reduces memory, exposes whole-scene context, and fits scarce-label regimes well.Partition quality caps accuracy; boundary recall, under-segmentation, over-segmentation, and partition-version evidence need their own sidecar.
Serialized transformer / foundation-model ceilingPoint Transformer V3, OctFormer, SphereFormerThe project has public or internal pre-training, enough GPU budget, and wants an accuracy-ceiling comparison after the conservative baselines are stable.Highest ceiling under good pre-training and tiling; serialized attention and octree/window variants carry more context than plain local convolution.Heavier training and inference; warmup, optimizer stability, tile halo, no-clipping boundary policy, and ensemble cost must be held constant before claiming a real gain.
Projection and standard-op deployment laneWaffleIron, SalsaNext, RangeFormer, FRNet, Cylinder3D, FlatFormerThe team needs a dependency-light baseline, an embedded-friendly single-scan sibling, or a fast projection model for color/intensity planes.Standard dense ops are easier to implement, optimize, and deploy; good for fast ablations and single-scan/map taxonomy alignment.Projection resolution discards some 3D detail; not a default release-map backbone unless thin-object, vertical-structure, and occlusion audits pass.
Image-distilled and open-vocabulary label lane2DPASS, OpenScene, Mosaic3D, LOSCCalibrated imagery, 2D foundation models, or open-vocabulary labels can improve supervised coverage, reviewer triage, or long-tail class discovery.Adds image supervision without necessarily adding an image dependency at release time; helps candidate labels for rare assets and semantic gaps.Candidate and pseudo labels are not release truth; projection QA, reviewer promotion, taxonomy mapping, and manifest permissions decide whether labels can be consumed.
SSM / Mamba efficiency frontierPoint-Cloud Mamba / SSM BackbonesA P3/P4 experiment needs larger context per tile or lower sequence cost after mature baselines define the target.Linear token cost is structurally attractive for registered maps where context radius and tile count dominate cost.Research-stage for release maps; ordering, rotation, density shift, local geometry recovery, seam stability, and non-road transfer need explicit audits before production use.

Minimum bake-off contract: compare the lanes above only under the same source-map acceptance package, taxonomy, release-state labels, tile manifest, halo policy, train/validation/test split, optimizer budget, augmentation schedule, and modality lane. Report semantic mIoU, class-balanced mIoU, rare/thin-class recall, boundary F1, seam disagreement, calibration error, throughput, cost per km2 or per 100M points, release-state confusion, false-permanent rate, and false-deletion rate. This prevents the common failure where a higher leaderboard mIoU hides an unusable runtime dependency, contaminated training export, or brittle non-road district transfer.

Expansion Backlog

The first waves focused on methods already identified as P0/P1 in the Perception Coverage Audit. The 2026-05-09 loops promoted SplatAD, GaussianFormer, GaussianOcc, streaming Gaussian occupancy, Cam4DOcc, StreamingFlow, Sparse4D, TacoDepth, RaCFormer, LIORNet, learned LiDAR desnowing/denoising, broad artifact removal, classical outlier filtering, MotionSeg3D, MambaMOS, neural scene-flow priors, CVFusion, 4D radar-camera occupancy, POD/FMCW LiDAR, DrivingGaussian, HUGS, SplatFlow, DistillNeRF, TrackOcc, cross-domain scene flow, LiDAR-camera occupancy fusion, dynamic occupancy/free-space, radar-LiDAR adverse-weather detection, RobuRCDet, SAMFusion, spatiotemporal memory occupancy flow, OVAD/OVODA, and open-vocabulary panoptic occupancy into atomic files. Future waves should split remaining grouped rows into atomic pages, especially:

  • VEON, ProOOD, and SA-Occ. EvOcc is already promoted; DR-REMOVER and ExelMap are routed through DR-REMOVER and ExelMap rather than duplicate perception pages.
  • Drive-OccWorld and DFIT-OccWorld where they need separate world-model or planning-facing treatment beyond the dynamic occupancy page.
  • DySS is routed through the sparse-query and streaming-temporal overviews pending official code or broader adoption; remaining sparse-query or end-to-end follow-ons should only become atomic pages when they add a non-duplicative method boundary beyond SparseBEV, DETR4D, Sparse4D, ForeSight, SparseDrive, or DiffusionDrive.
  • LinkOcc and related 2025-2026 occupancy follow-ons; GaussianFlowOcc, GaussTR, and GS-Occ3D are now first-class pages, while missing-view occupancy routes through M2-Occ, temporal radar-camera occupancy routes through TEOcc, and Gaussian-rendered occupancy routes through GaussRender unless a distinct method needs its own page.
  • CoDS, JigsawComm, and collaborative Gaussian follow-ons such as GSCOOP. VOGS-CP is already promoted as the collaborative Gaussian semantic-occupancy reference, SparseCoop is already promoted as the sparse cooperative-query reference, QuantV2X is already promoted as the cooperative quantization reference, and TruckV2X is promoted as the truck-centered cooperative dataset reference.
  • Residual FOD OOD-vs-detector benchmark planning and airside-specific dust/de-icing-mist/steam/glycol/wet-apron datasets; SpaCeFormer now covers fast proposal-free open-vocabulary 3D instance segmentation, EmbodiedScan/MMScan covers embodied robotics 3D perception benchmarks, and DriveBench stays discoverable through the VLA/VLM reliability benchmark page rather than duplicating it here.

Sources

Public research notes collected from public sources.