ACM MM 2026 ยท Dataset Track

TransBiolab

A Real-World Multi-View Dataset of Cluttered Transparent Biomedical Objects

Ke Ma1,2 Yifei Wang3 Meng Wang2,4 Tian Xia3โ€ 
1School of AI & Automation, HUST 2College of Design & Innovation, Tongji University 3School of Software and Engineering, HUST 4Shanghai Institute for Intelligent Autonomous Systems, Tongji โ€ Corresponding author

TransBiolab is a real-world RGB-D dataset of cluttered transparent biomedical objects captured as calibrated multi-view sequences. It provides 6D poses, full and visible masks, depth, per-frame camera calibration, and 15 OBJ-format CAD models โ€” enabling segmentation, depth estimation, 6D pose estimation, and multi-view reasoning for autonomous laboratory manipulation.

15
Object Types
98
Scenes
161K
RGB-D Frames
1.03M
Annotations
11
Arrangements
4
Backgrounds
2
Lighting

Objects & Capture Setup

15 laboratory objects organized into five functional groups that frequently appear in biomedical manipulation workflows.

TransBiolab overview teaser
An overview of the proposed TransBiolab dataset, featuring multi-object and multi-view characteristics. Frames denoted by red rectangles indicate higher levels of clutter and occlusion.

TransBiolab contains objects from five groups: well plates (6-, 12-, 24-, 96-well), cell culture dishes (35mm, 60mm, 90mm), liquid-handling objects (15ml and 50ml centrifuge tubes, pipette, tube rack), cell culture flasks (25ml, 75ml, 125ml), and a 1L bioreactor.

Seven objects are fully transparent, six combine transparent bodies with opaque caps or bases, and two are opaque auxiliary tools. This composition reflects how transparent vessels coexist with supporting tools in real laboratory workcells.

Data are collected with an Intel RealSense D435i RGB-D camera at 1280ร—720 resolution and 30 FPS, mounted on a 7-DoF Franka Emika Panda robot arm that follows calibrated multi-view trajectories around the scene.

15 TransBiolab objects
Fig. 1: Rendering views of the 15 objects included in TransBiolab.
Well Plates
5 6-well plate
6 12-well plate
7 24-well plate
18 96-well plate
Cell Culture Dishes
9 35mm dish
3 60mm dish
11 90mm dish
Liquid Handling
13 15ml centrifuge tube
14 50ml centrifuge tube
1 Pipette
2 Tube-rack
Flasks & Bioreactor
15 25ml flask
16 75ml flask
17 125ml flask
19 Grex 1L bioreactor

Scene Construction & Capture

Calibrated Multi-View Trajectory

Each scene is captured as an ordered multi-view sequence by the robot-mounted camera following a calibrated hemispheric path. This provides per-frame camera intrinsics and extrinsics, supporting both discrete-view evaluation and future multi-view or video-based methods.

For benchmarking, we uniformly sample 10 equally spaced viewpoints from each trajectory, covering elevation angles from -22ยฐ to -62ยฐ and varying azimuth angles.

Robot-mounted capture trajectory
Fig. 2: Robot-mounted capture trajectory with camera pose visualization.

11 Arrangement Patterns

Scenes are constructed from 11 arrangement patterns designed around real laboratory procedures: sample preparation, liquid handling, mixing, sampling, transport, and cell expansion. Patterns range from moderate clutter (4โ€“5 objects) to dense scenes (9โ€“12 objects) with heavy mutual occlusion.

Some patterns emphasize repeated instances from the same category (e.g., multiple tubes), while others mix several functional groups and include opaque auxiliary tools.

11 arrangement patterns
Fig. 3: Eleven arrangement patterns used to build controlled scenes.

Background & Lighting Variations

Controlled scenes are varied by four tabletop backgrounds (white tablecloth, grey tablecloth, bare wood, magazine-covered) and two lighting settings (top light ~190 lx, side light ~75 lx). These factors change contrast, reflections, and occlusion cues without altering scene composition.

Background and lighting variations
Fig. 4: Four backgrounds and two lighting settings.
Scene-object distribution matrix
Fig. 6: Scene-object distribution matrix and initial scene-disjoint training/test partition. โ† Scroll horizontally to view full detail โ†’
Dataset metadata
Fig. 8a: Compact metadata visualization โ€” objectwise viewpoint coverage and the released 6D annotation distributions over translation and rotation axes.
6D annotation distributions
Fig. 8b: 6D pose annotation distributions across translation and rotation axes.

๐Ÿ“ Scene Directory Structure

scene_XXXXXX/
โ”œโ”€โ”€ rgb/                             # 1280ร—720 color images
โ”‚   โ”œโ”€โ”€ 000000.png  โ€ฆ  NNNNNN.png
โ”œโ”€โ”€ depth/                           # 16-bit depth maps (mm)
โ”‚   โ””โ”€โ”€ โ€ฆ
โ”œโ”€โ”€ mask/                            # Full projected masks (incl. occluded)
โ”‚   โ”œโ”€โ”€ {frame}_{obj_idx}.png
โ”œโ”€โ”€ mask_visib/                      # Visible-only masks
โ”‚   โ””โ”€โ”€ โ€ฆ
โ”œโ”€โ”€ scene_camera.json                # Per-frame K matrix + depth_scale
โ”œโ”€โ”€ scene_gt.json                    # 6D pose (R, t, obj_id) per instance
โ””โ”€โ”€ scene_gt_info.json               # Visibility fraction + bbox

Benchmark Results

Evaluation protocols for FoundationPose (6D pose estimation) and SAM 3 (open-set segmentation), stratified by object category, scene clutter, and viewpoint.

4.1  FoundationPose for 6D Pose Estimation

For each target instance, we load the RGB-D frame, visible mask, camera intrinsics, and CAD model, then run the standard register-refine pipeline with 15 refinement iterations. We report ADD, ADD-S, and AUC up to 0.1m.

FoundationPose
6D Pose Estimation
80.8% ADD-S AUC
ADD AUC53.9%
Mean ADD-S23.2 mm
Instances Evaluated1,892
SAM 3
Open-Set Segmentation
68.5% mIoU
Match Rate84.8%
Matched mF10.881
Instances Evaluated1,247

Table 2: FoundationPose Results by Object Category

NameADD-S AUCADD AUCADD-S (mm)
Pipette93.088.210.8
Tube-rack94.146.29.9
60mm dish90.156.99.9
6-well plate91.749.28.6
12-well plate91.944.28.1
24-well plate93.452.86.6
35mm dish83.866.717.0
90mm dish90.637.611.0
15ml centrifuge tube75.464.424.8
50ml centrifuge tube81.058.719.7
25ml flask77.757.838.1
75ml flask54.235.053.2
125ml flask41.028.175.1
96-well plate91.653.58.6
Grex 1L bioreactor72.945.431.9
Overall80.853.923.2

Table 3: FoundationPose by Geometric Family

ShapeRepresentative ObjectsADD-S AUCADD-S (mm)
Plate-like6/12/24/96-well plate92.27.9
Rack / elongatedPipette, tube-rack93.610.4
Dish-like35/60/90mm dish90.410.4
Tube-like15/50ml centrifuge tube78.222.2
Bottle-like25/75ml flask65.945.6
Large / irregular125ml flask, Grex 1L bioreactor56.953.5

Table 4: FoundationPose Across 10 Sampled Viewpoints

ViewElev.Azim.ADD-S AUCADD AUCADD-S (mm)
VP1-22.2ยฐ-68.1ยฐ72.247.238.7
VP2-24.6ยฐ-99.0ยฐ72.243.248.4
VP3-27.3ยฐ-74.6ยฐ74.041.628.1
VP4-32.6ยฐ-86.1ยฐ73.948.730.1
VP5-40.9ยฐ-85.0ยฐ79.954.121.3
VP6-51.1ยฐ-86.2ยฐ88.060.312.2
VP7-51.9ยฐ-103.1ยฐ84.756.316.1
VP8-53.4ยฐ-75.2ยฐ87.160.313.3
VP9-58.9ยฐ-83.4ยฐ87.664.812.5
VP10-62.0ยฐ-75.6ยฐ87.861.512.5
๐Ÿ’ก
Viewpoint effect: Steep views (>45ยฐ) achieve 85โ€“88% ADD-S AUC (mean 12.5mm), while shallow views (<33ยฐ) only reach 72โ€“74% (mean 30โ€“48mm). Mid-to-steep viewpoints expose larger top surfaces and clearer contours of transparent containers.

Scene Complexity Analysis

FoundationPose results by scene complexity
Fig. 9: FoundationPose ADD-S AUC grouped by total visible objects (top) and distinct categories (bottom). Performance is dominated by object geometry rather than scene density.

4.2  SAM 3 for Open-Set Segmentation

SAM 3 is queried with text prompts per object category. Predicted masks are greedily matched to GT instances by IoU. Unmatched GT masks receive IoU = 0.

Table 5: SAM 3 Results by Object Category

NameMatch RatemIoUmF1
Pipette89.9%0.7180.884
Tube-rack90.0%0.6400.828
60mm dish95.5%0.8180.922
6-well plate85.0%0.7660.946
12-well plate87.8%0.7970.951
24-well plate87.8%0.8020.953
35mm dish83.3%0.7200.900
90mm dish97.1%0.8450.925
15ml centrifuge tube71.7%0.4840.795
50ml centrifuge tube75.6%0.4950.780
25ml flask91.3%0.7350.883
75ml flask77.1%0.5740.829
125ml flask78.6%0.5870.799
96-well plate81.4%0.7210.932
Grex 1L bioreactor100.0%0.9140.954
Overall84.8%0.6850.881

Table 6: SAM 3 by Geometric Family

ShapeRepresentative ObjectsmIoUMatch Rate
Plate-like6/12/24/96-well plate & 35/60/90mm dish0.78989.2%
Bottle-like25/75/125ml flask, tube-rack, Grex 1L0.70086.3%
Tube-like15/50ml centrifuge tube & pipette0.53676.1%

Table 7: SAM 3 Across 10 Sampled Viewpoints

ViewElev.Azim.mIoUMatch RatemF1
VP1-22.2ยฐ-68.1ยฐ0.68188.0%0.856
VP2-24.6ยฐ-99.0ยฐ0.68191.1%0.837
VP3-27.3ยฐ-74.6ยฐ0.68489.6%0.847
VP4-32.6ยฐ-86.1ยฐ0.68586.3%0.874
VP5-40.9ยฐ-85.0ยฐ0.68785.5%0.880
VP6-51.1ยฐ-86.2ยฐ0.69382.4%0.905
VP7-51.9ยฐ-103.1ยฐ0.68284.0%0.886
VP8-53.4ยฐ-75.2ยฐ0.69282.4%0.905
VP9-58.9ยฐ-83.4ยฐ0.60971.2%0.919
VP10-62.0ยฐ-75.6ยฐ0.58368.0%0.915
๐Ÿ’ก
SAM 3 viewpoint pattern: From VP1 to VP8, mIoU remains stable (0.681โ€“0.693). The main degradation appears at steepest views โ€” VP9 and VP10 drop to 0.609 and 0.583 mIoU, with match rates falling to 71.2% and 68.0%. Failure at steep views is driven by missed detections rather than poor mask quality (mF1 stays high at 0.919/0.915).

Additional Baselines & Analyses

We thank all reviewers for the constructive comments. We address the main concerns below. TBiolab-HO denotes the 10 held-out lab scenes.

P1. Additional Baselines

To Reviewers zrnj, 5Exs, and TgfS, we added TransLab, ClearGrasp, DA3, and MegaPose-6D, while keeping SAM3 and FoundationPose.

Segmentation

Dataset-defined masks are used; this mask-level protocol complements the paper's instance-matching SAM3 protocol. All metrics are higher-is-better.

DatasetMethodTargetIoUPrecisionRecallF1
TransBiolabSAM3object mask0.6620.8810.7030.764
TransBiolabTransLabobject mask0.3630.5860.6530.496
ClearPoseSAM3object mask0.6990.8850.7440.789
ClearPoseTransLabobject mask0.4910.5690.8530.632
Trans10KSAM3object mask0.7040.8210.8270.824
Trans10KTransLabobject mask0.8120.8740.9160.881
TBiolab-HOSAM3object mask0.4770.6560.6480.632
TBiolab-HOTransLabobject mask0.3440.4750.6650.491
โ†˜
This verifies domain specificity: TransLab excels on Trans10K but drops on TransBiolab and TBiolab-HO; SAM3 is lower on TransBiolab than ClearPose, so labware clutter remains unsaturated.

Depth

DA3 uses metric depth, no test-time scale or shift alignment. All metrics are lower-is-better.

DatasetSplitMethodReferenceAbsRelRMSE (m)MAE (m)Raw missing
TransBiolaballDA3depth GT0.3710.2240.2200.349
TransBiolaballClearGraspdepth GT0.3930.4730.3490.313
ClearPoseallDA3depth GT0.1610.2210.1410.390
ClearPoseallClearGraspdepth GT0.3270.4040.2760.299
TransBiolabheld-outDA3depth GT0.3020.1800.1740.189
TransBiolabheld-outClearGraspdepth GT0.1790.1320.0860.178
โ†˜
Compared with ClearPose-all, both DA3 and ClearGrasp have higher object-region depth errors on TransBiolab-all, suggesting that TransBiolab remains challenging on depth estimation and completion tasks.

6D Pose

FoundationPose uses RGB-D, GT mask, CAD; MegaPose-6D uses RGB, GT bbox, CAD, no depth. AUC and Success are higher-is-better; Mean error is lower-is-better.

DatasetSplitMethodInputADD-S AUCADD AUCADD-S MeanSuccess
TransBiolaballFoundationPoseRGB-D, GT mask, CAD80.853.923.2 mm67.86%
ClearPoseallFoundationPoseRGB-D, GT mask, CAD91.0549.269.77 mm89.53%
TransBiolaballMegaPose-6DRGB, GT bbox, CAD, no depth31.7631.4967.76 mm34.14%
ClearPoseallMegaPose-6DRGB, GT bbox, CAD, no depth39.8412.4064.29 mm25.00%
TransBiolabheld-outFoundationPoseRGB-D, GT mask, CAD75.3547.0634.53 mm62.31%
TransBiolabheld-outMegaPose-6DRGB, GT bbox, CAD, no depth55.5025.1796.46 mm26.59%
โ†˜
This verifies pose remains hard beyond depth artifacts: FoundationPose drops on TransBiolab and TBiolab-HO, while depth-free MegaPose-6D remains low, showing domain and geometry gaps.

P2. Annotation Reliability

To Reviewers W5GB, TgfS, and 5Exs, we added a geometry-consistency audit over CAD, poses, masks, intrinsics, and trajectories; it measures mutual consistency, not independent metrology.

Pose-mask consistencyvisible-mask IoU: mean 0.949; median 0.971
Same-view reprojectionmedian symmetric contour error: 0.62 px
Cross-view consistencytransferred median contour error: 0.91 px

The 95th-percentile contour error is below 1 px for both audits. We will add overlays and the protocol.

P3. Held-out, Stats, and Redundancy

To Reviewers jqnR, 5Exs, TgfS, and W5GB, held-out scenes above have new layouts, more distractors, and different backgrounds and lighting.

Release size98 scenes, 161,315 frames
Typical lengthabout 1,646 frames per scene
Scene setupeach scene has more than 3 objects
>3 in-frame objects161,059 frames, 99.84%
exactly 3 in-frame objects256 frames, 0.16%

Exactly-3-object frames occur briefly when viewpoint motion moves objects out of frame; we retain them for continuity.

P4. Robot Experiment

To Reviewers W5GB, 5Exs, and TgfS, this is system-level manipulation, not a proxy for 6D pose accuracy. Parallel-jaw grasping is low-dimensional and error-tolerant, while LinkerHand needs affordance and grasp-type selection, high-dimensional joint mapping, and contact-rich closure.

Its lower success reflects grasp synthesis and control complexity, not only pose error.

P5. Minor Revisions

To Reviewers jqnR, 5Exs, and TgfS, we will fix captions, crop Fig. 1, enlarge figure and table fonts, and reformat tables.

We will clarify that 10 viewpoints are diagnostics while full trajectories are released, and 1,300-1,800 frames follows about 50s 30FPS capture with timing variation.

Visualization Comparison

Side-by-side SAM 3 segmentation masks and FoundationPose 3D bounding box projections on complex multi-object scenes.

1 / 20 Scene 000013 ยท Frame 000102 ยท VP01
Visualization

Real Robot Experiment

The system first obtains an object mask from segmentation, then estimates a target 6D pose with FoundationPose, transforms this pose to the robot base through hand-eye calibration, and finally plans grasp and placement motions through ROS/MoveIt.

๐Ÿฆพ Franka Parallel-Jaw
65.3% 98 / 150 trials
๐Ÿค– LinkerHand 10-DoF
56.67% 85 / 150 trials
Robot platform
Fig. 10: Real robot platform used for cluttered transparent-object manipulation experiments.

๐ŸŽฌ Robot Manipulation Demo

Key Findings

01

Plate & Dish Objects Excel

Flat plate-shaped objects achieve 90โ€“94% ADD-S AUC and 0.789 mIoU. Their stable contours benefit both pose estimation and segmentation.

02

Large Transparent Bottles Struggle

75ml and 125ml flasks score as low as 41โ€“54% ADD-S AUC. Large transparent volumes cause severe depth and feature ambiguity.

03

Viewpoint Matters for Pose

FoundationPose gains +15.6% AUC from shallow (22ยฐ) to steep (62ยฐ) views. SAM 3 shows the opposite: degradation at steepest views due to missed detections.

04

Tube-like Objects Are Hardest for SAM 3

Tubes and pipettes achieve only 0.536 mIoU with 76.1% match rate. Thin cylindrical shapes with weak boundaries challenge text-prompted segmentation.

05

Symmetry Exposes ADD vs ADD-S Gap

For symmetric objects the 26.9% gap between ADD-S and ADD confirms that transparent objects need symmetry-aware metrics.

06

Clutter Interacts with Geometry

Scene density alone does not predict difficulty. Performance depends on object geometry ร— symmetry ร— visibility, exactly what TransBiolab is designed to surface.