How e-con Systems enables Reliable Low-Level Obstacle Detection in AMRs with a 3D CW-iToF Camera – DepthVista Helix

It’s common knowledge that Autonomous Mobile Robots (AMRs) usually navigate confidently until something small gets in the way. For instance, pallets, pipes, toolboxes, and loose packages can be found below the scanning plane of a 2D LiDAR yet remain invisible to the navigation stack until the robot reaches them.

e-con Systems integrated the DepthVista Helix, an iToF camera, into a ROS 2 Humble rover running Nav2 on the CV72eSOM, turning dense depth data into a 3D voxel obstacle layer the planner can act on. In the video below, you’ll see that the same rover runs the same route twice – once with DepthVista Helix active, and once with the 2D LiDAR alone.

The video explains how the DepthVista Helix helps low-lying objects enter the costmap early enough for Nav2 to plan around them.  With the 2D LiDAR alone, objects below the scan plane may not appear in the scan.

In this blog, you’ll learn how the perception pipeline works, why each design decision was made, and which problems we had to solve to make it run reliably on embedded hardware.

Importance of Low-Level Obstacle Detection

Autonomous Mobile Robots (AMRs) are transforming warehousing, manufacturing, logistics, and healthcare by automating material handling, inspection, and transportation. Navigation stacks have become increasingly sophisticated, yet reliable low-level obstacle detection remains one of the hardest problems to solve in real deployments.

A 2D LiDAR observes a single horizontal slice of the world. Anything below that slice produces no return at all, so pallets, pipes, toolboxes, cables, and small packages don’t exist as far as the navigation stack is concerned. A single undetected obstacle can lead to a collisions, equipment damage, operational downtime, and potential safety risks.

DepthVista Helix: e-con Systems’ High-Performance iToF Camera

e-con Systems’ DepthVista Helix is a 3D CW indirect Time-of-Flight (iToF) camera that measures the distance to every pixel in the scene rather than just its intensity, producing a dense depth map that describes obstacles by both position and height.

To validate the approach, the camera was integrated into an autonomous rover running ROS 2 Humble and the Nav2 navigation stack alongside an RPLIDAR S2E, with the complete perception pipeline running on the e-con Systems CV72eSOM. The entire stack operates on embedded hardware, with no off-board compute.

Let’s now understand the practical benefits of low-level obstacle detection.

What Are the Practical Benefits of Low-Level Obstacle Detection?

Detection below the LiDAR scan plane

The camera covers exactly the volume the 2D LiDAR can’t see, from the floor up to the scan plane. Obstacles that previously caused unexplained collisions become visible to the planner well before the robot reaches them.

Height-aware obstacle representation

Since every point carries a height, the costmap can distinguish a low curb from a toolbox and floor noise from a real object. Obstacles are described by volume rather than by a single flat outline.

Stable navigation across different floor finishes

Polished concrete, epoxy coatings, and reflective tiles all generate different levels of depth noise. Filtering and layer tuning keep the obstacle map consistent as the robot moves between areas with different surfaces.

Embedded-ready on resource-constrained hardware

Depth acquisition, point cloud generation, 3D obstacle mapping, localization, and path planning all run together on the CV72eSOM. No external GPU or cloud connectivity is required, which keeps the platform suitable for industrial and field deployment.

Reusable 3D maps across a robot fleet

The dense perception data can be used to build static maps for ROS 2 navigation. Once created, a map can be reused by the same robot or deployed across multiple robots operating in the same environment.

What’s Inside e-con Systems’ Low-Level Obstacle Detection & Navigation Solution?

The solution is built around five core components that work together inside the ROS 2 ecosystem:

  • Depth image acquisition with configurable publishing and cropping
  • Point cloud generation from iToF depth frames
  • LiDAR-camera sensor fusion within the navigation costmap
  • 3D obstacle mapping using the Spatio-Temporal Voxel Layer
  • Nav2 path planning and waypoint navigation

Each component runs as an independent ROS 2 node communicating through standardized topics and services, which keeps the perception pipeline modular, configurable, and easy to replace piece by piece.

The key technologies powering this system are:

  • ROS 2 Humble
  • Nav2 Navigation Stack with AMCL Localization
  • DepthVista Helix
  • RPLIDAR S2E
  • Spatio-Temporal Voxel Layer (STVL)
  • e-con Systems CV72eSOM (Ambarella CV72)
  • Custom ROS 2 C++ and Python nodes

ROS 2 Low-Level Obstacle Detection: Understanding the Workflow

The pipeline runs in four stages, from the raw depth frame to the costmap the Nav2 planner consumes. Every stage was tuned around describing the floor-adjacent volume in front of the robot as accurately as possible, at the lowest processing cost.

Capturing depth and choosing the operating mode

e-con Systems’ DepthVista Helix captures high-density depth images, and each frame is converted into a 3D point cloud in which every valid pixel becomes a point with X, Y, and Z coordinates. This cloud feeds everything downstream, so the choice of operating mode shapes the quality of the final obstacle map.

HD Mode VGA Mode
Resolution 1280 x 960 640 x 480
Depth range 0.2 – 2 m 0.2 – 6 m
Operation Single frequency Dual frequency
Optimized for High-resolution depth sensing Real-time navigation

VGA mode was selected, and the reasoning was range and accuracy rather than resolution:

  • Range determines reaction distance. A 2 m limit leaves a moving rover almost no room, so obstacles arrive as hard stops. A 6 m range lets them enter the costmap early enough for the planner to route around them.
  • Dual frequency reduces measurement error. Two modulation frequencies suppress phase ambiguity and lower depth noise, so fewer spurious points reach the voxel grid, and fewer false obstacles reach the planner.
  • Resolution is not the limiting factor. Once the cloud is discretized into voxels, the extra pixels of HD mode collapse into the same cells – added CPU load, no added obstacle information. VGA mode also returns more stable depth from dark, low-reflectivity surfaces at range.

Cropping and aligning the camera to the region of interest (ROI)

The RPLIDAR S2E already covers the horizontal scan plane, so the camera only has to describe the volume between the floor and that plane, directly in front of the robot. Three things had to be right.

  • Mounting and tilt. The camera sits low on the front of the chassis, tilted down so the sensing cone starts just ahead of the bumper. Mounting it higher opens a blind zone exactly where an undetected obstacle is most dangerous.
  • Transform alignment. The transform between the robot base and the camera optical frame is measured. At a shallow downward angle, a fraction of a degree of pitch error either lifts the entire floor into the costmap or drops real obstacles below the height threshold.
  • Depth image cropping. Only a horizontal band of the depth image matters. Cropping to that band before point cloud generation sharply cuts the points produced per frame, with no loss of information the application uses.

Getting placement, calibration, and cropping right here removes more noise than any amount of filtering further down the pipeline.

Building the 3D obstacle map

The point cloud is converted into a 3D occupancy grid by dividing the space around the robot into small volumetric cells called voxels, so an obstacle is described by its height as well as its footprint. ROS 2 offers two layers for this, and the choice mattered more than expected.

The standard Voxel Layer clears voxels by ray tracing. A voxel is freed only when a sensor ray is observed passing through it. With a downward-tilted camera, floor-adjacent voxels are rarely re-observed once the robot has moved past them, so stale obstacles linger and block paths that are actually free. The Spatio-Temporal Voxel Layer (STVL) instead gives every voxel a timestamp and expires it after a configurable period, so clearing never depends on line of sight.

Aspect Voxel Layer Spatio-Temporal Voxel Layer
Data structure Dense, fixed-size 3D grid Sparse voxel grid
Clearing method Ray tracing through free space Time-based decay
Memory usage Scales with the mapped volume Scales with occupied voxels
Stale obstacles Persist until ray-traced clear Expire after the decay period
Embedded CPU cost Higher Lower

The layer was then tuned around one objective, describing a thin slab of space in front of the robot rather than the full sensing volume:

  • Voxel size matched to the smallest obstacle the rover must avoid – fine enough to resolve it, coarse enough to keep memory and update cost down.
  • Height limits set just above the floor plane and capped near the LiDAR scan plane, so floor noise is discarded and obstacles are never counted twice.
  • Decay time and minimum points per voxel tuned so obstacles persist long enough for the planner to route around them, while isolated speckle returns from dark surfaces are rejected.

The result is a voxel map that is dense where real obstacles are and empty everywhere else, which is exactly what the planner needs to generate smooth, collision-free paths.

Generating the navigation costmap

The obstacle information from the STVL is merged with the LiDAR data into the robot’s navigation costmap. The costmap assigns a traversal cost to every region, marking occupied space as high-cost and identifying the free space around it. The Nav2 planner continuously uses this map to compute collision-free paths, so a low-profile obstacle detected by the camera influences the route in exactly the same way as a wall detected by the LiDAR.

How e-con Systems Solved Key Obstacle Detection Challenges

Deploying an iToF camera on an autonomous robot involves far more than connecting a sensor. The perception pipeline has to deliver accurate depth while sharing a single embedded processor with localization, navigation, and control. Most of the engineering effort went into these four problems.

1) Running high-bandwidth depth perception on a single embedded device

The challenge

Continuously processing and publishing high-bandwidth depth data can saturate the CPU on an embedded platform, leaving insufficient headroom for critical ROS 2 components such as localization, navigation, and path planning. The camera had to earn its place without starving the rest of the autonomy stack.

Our solution

The image publishing pipeline was optimized through configurable publish rates, selective topic publishing, configurable image cropping, and further processing improvements. These changes reduced unnecessary computation and lowered the CPU utilization of both the depth image publishing and the point cloud generation pipelines while preserving the quality of the perception data.

Business impact
  • Perception, navigation, and control run together on the CV72eSOM
  • No external GPU or off-board compute is required in the field
  • More processing headroom remains for additional perception algorithms

2) Separating genuine obstacles from floor noise

The challenge

Different floor finishes produce different levels of depth noise. While capturing the depth of low-level obstacles, the camera also captured returns from the floor itself, which entered the costmap as false obstacles and caused unstable navigation. A pipeline tuned for one surface could behave poorly in the next aisle.

Our solution

The perception pipeline was tuned end-to-end rather than at a single point. Confidence thresholding and cropping at the camera, a minimum obstacle height set just above the floor plane, and a minimum points-per-voxel threshold in the STVL together removed floor noise while preserving genuine obstacles.

Business impact
  • Reliable obstacle mapping across different indoor environments
  • Fewer false detections and fewer unnecessary stops
  • More stable and predictable autonomous navigation

3) Detecting dark and low-reflectivity objects

The challenge

Reliable detection depends on obstacle height, material, distance, and camera placement. Dark or low-texture objects return fewer valid depth points, which makes them harder to detect consistently and easy to lose at the far end of the range.

Our solution

The camera was characterized across a range of real-world test scenarios and validation samples covering different materials, obstacle heights, and distances. Mounting position, tilt angle, and perception parameters were then optimized around those results, and VGA mode with dual-frequency operation was selected for its more stable returns from poorly reflecting surfaces.

Business impact
  • Reliable detection of low-profile obstacles in real conditions
  • Validated operating limits that deployment teams can design against
  • Improved navigation safety in cluttered environments

4) Minimizing phase wrapping artifacts

The challenge

Phase wrapping is an unavoidable physical phenomenon in indirect Time-of-Flight cameras. Objects beyond the measurable range appear at incorrect, much closer distances, which injects false obstacles directly into the navigation pipeline.

Our solution

Camera parameters, including the confidence threshold and the configured sensing range, were carefully tuned to suppress ambiguous returns, and dual-frequency operation was used to resolve the ambiguity where possible. Range limiting in the point cloud stage discards the remaining out-of-range returns before they reach the costmap.

Business impact
  • More accurate depth measurements across the working range
  • Reduced false obstacle detection at long distances
  • Greater confidence in autonomous navigation decisions

How e-con Systems’ iToF Camera Elevates AMR Navigation

Reliable low-level obstacle detection matters more as autonomous robots move into complex industrial and commercial environments. Traditional perception stacks struggle with anything below the LiDAR scan plane, and the DepthVista Helix, an iToF camera, closes exactly that gap by supplying the dense 3D depth information a planner needs to route around a pallet instead of into it.

What made the difference in practice were the decisions around it. This included selecting VGA mode for range and accuracy over raw resolution, mounting and calibrating the camera around a single well-defined region of interest, cropping the depth image before it becomes a point cloud, and choosing a voxel layer whose clearing behavior matches how a downward-tilted camera actually observes the world.

The result is an autonomy stack that detects low-profile obstacles without sacrificing real-time performance, running entirely on embedded hardware. Whether deployed in warehouses, manufacturing facilities, hospitals, or service robotics, it offers a practical and scalable way to improve robotic perception and navigation safety.

The complete perception pipeline running live — cropped depth frames on the left, the Nav2 costmap built from them in the center, and the real scene on the right. The white points scattered through the costmap are occupied voxels from the STVL, marking the low-profile obstacles the 2D LiDAR can’t see.

Want to Learn More About Low-Level Obstacle Detection?

If you would like to know more about this work or explore how the technology could apply to your use case, please fill out the form below and our team will be in touch.

https://www.e-consystems.com/Request-form.asp

e-con Systems: A Leader of Embedded Vision Innovation

e-con Systems has been designing, developing, and manufacturing embedded vision solutions, from OEM cameras to complete ODM platforms, since 2003.

Use our Camera Selector to identify the vision solution that matches your requirements.

Need help choosing the right camera? Connect with our vision experts at camerasolutions@e-consystems.com to better understand the selection process.

FAQs

What is low-level obstacle detection in autonomous mobile robots?

It is the ability to detect objects that sit below the scanning plane of a 2D LiDAR, such as pallets, pipes, cables, and small packages, and to represent them in the navigation costmap so the robot can plan around them.

Why is a 2D LiDAR not enough on its own?

A 2D LiDAR observes a single horizontal slice of the environment. Anything below that slice produces no return, so the navigation stack has no information about it until a collision occurs.

How does the iToF camera data reach the Nav2 planner?

Each depth frame is converted into a 3D point cloud, inserted into the Spatio-Temporal Voxel Layer as occupied voxels, and merged into the navigation costmap that the Nav2 planner uses to compute collision-free paths.

Why was the Spatio-Temporal Voxel Layer chosen over the standard Voxel Layer?

STVL clears voxels using a time-based decay model instead of ray tracing. With a downward-tilted camera that rarely re-observes floor-adjacent voxels, that behavior removes stale obstacles predictably and uses less memory and CPU on embedded hardware.

Why use VGA mode instead of the higher-resolution HD mode?

VGA mode extends the depth range from 2 m to 6 m and uses dual-frequency operation to reduce depth noise. Since the point cloud is discretized into voxels anyway, the extra HD resolution adds processing load without improving the obstacle map.

Which technologies power the solution?

ROS 2 Humble, the Nav2 navigation stack with AMCL localization, the DepthVista Helix, an RPLIDAR S2E, the Spatio-Temporal Voxel Layer, the e-con Systems CV72eSOM, and custom ROS 2 C++ and Python nodes.

Related posts

Introducing NetraVis, e-con Systems’ 1MP Active Stereo Depth Camera for Accurate Robotic Perception

How e-con Systems Solves Motion and Lighting Challenges in High-Speed Imaging

Edge AI vs. Cloud ANPR: Where Should the Decision Happen for Traffic Enforcement?