3D Dataset Visualization Overview

What do you see?
Motivation
Purpose of visualization
- To understand - comprehend a dataset in all spatial dimensions at once.
- To learn - visualize specific features to draw conclusions from.
- To share and tell - discuss your work (and your data) with and beyond scientific circles.
Motivation
From 2D to 3D: Adding depth
- Occlusion
- Shading and shadows
- Perspective
- Motion parallax
- Stereo / VR
3D Datasets - Data Types
Which axes does your data have?
- X, Y, Z - the three spatial axes. This is the part that makes it 3D. Also called width, height, and depth.
- C - channels: stains, wavelengths, labels.
- T - time: the same volume recorded at different time points.
- 3D in this workshop refers to the three spatial axes.
3D Datasets - Data Types
Voxels
Grid-based data structure: Voxels are values at every position in a discrete block
This enables us to look inside any structure in the grid
Watch for resolution, and whether it is isotropic or anisotropic
Formats TIFF / OME-TIFF, OME-Zarr, HDF5 / N5, NIfTI, NRRD, DICOM, MRC, vendor formats (CZI, LIF, ND2, IMS, LSM), NetCDF, …
3D Datasets - Data Types
Meshes
Points in space joined into triangles, describing a surface
Quality metrics:
- watertight: no holes, so the surface encloses a volume
- manifold: every edge is shared by exactly two triangles
- normals facing outward
Formats STL, PLY, OBJ, glTF / GLB, VTK / VTP
3D Datasets - Data Types
Point Clouds
Carries positions plus attributes; no connectivity
Drawn as primitives: a dot , **square, or a splat (a soft oriented disc that blends with its neighbours into a surface)
The density of the points will determine how well one can estimate a surface
Formats LAS / LAZ, COPC, E57, PLY, plain XYZ / CSV
3D Datasets - Data Types
Vector Fields
Carries a direction and a magnitude per location
Can be stored like voxels as a multi-channel volume
The magnitude can for example represent flow, displacement, or diffusion
Glyphs show directions for individual points; streamlines show where the field leads
Formats VTK / VTI / VTU, NetCDF, multi-component NIfTI, OME-Zarr or HDF5 with a component axis
3D Datasets - Sources in Science
Microscopy
- Examples: confocal, light-sheet, electron microscopy, histology
3D Datasets - Sources in Science
Tomography
- Examples: clinical CT, micro-CT, synchrotron tomography, electron tomography, MRI
3D Datasets - Sources in Science
Photogrammetry
- Examples: drone survey, handheld and phone capture, multi-camera rigs
3D Datasets - Sources in Science
Range scanning
- Examples: terrestrial laser scanning, airborne LiDAR, structured light, time-of-flight cameras
3D Datasets - Sources in Science
Echo and wave methods
- Examples: seismic reflection, sub-bottom profiling, sonar, ground-penetrating radar, ultrasound, imaging radar (SAR)
3D Datasets - Sources in Science
Simulation
- Examples: CFD, molecular dynamics, finite elements, climate models
Rendering Pipeline
How do we get from 3D data to a picture on a screen?
- The graphics pipeline, on learnopengl.com
- Akenine-Möller, Haines & Hoffman, Real-Time Rendering - chapter 2, the four-stage model used above
Rendering Pipeline
Terms which are helpful to know
Some terms come up repeatedly in the context of 3D rendering:
- Shader - a small program that runs on the GPU, once per vertex or once per pixel.
- Rasterization - for each triangle, which pixels does it cover? The step that maps freely positioned 3D objects onto a raster.
- Ray casting - for each pixel, send a ray into the scene and read what it passes through. Volume rendering is this.
- Ray tracing - ray casting plus secondary rays, which is where reflections and shadows come from.
How to pick a visualization approach
Questions to ask
- Explorative or reproducible - clicking, or scripting
- How accessible - what does your reader have to download and install
- Fast or pretty - realtime vs. fancy, expensive rendering
How to pick a visualization approach
Step 1 - what exactly is the artifact?
How to pick a visualization approach
Step 2 - what has to be visible, and what carries it?
How to pick a visualization approach
Step 3 - does it fit in memory?
- Fits comfortably - load it and work
- Fits, but barely - downsample to explore, run the real thing headless
- Does not fit - change the data layout:
- Chunking / tiling: splitting the dataset up into sections (TIFF, ZARR, HDF5, load via Dask)
- Resolution pyramids: storing multiple resolutions of different sizes (OME-ZARR)
- Load data on demand (Neuroglancer, BigDataViewer, BigVolumeViewer, …)
How to pick a visualization approach
Step 4 - choosing a tool - desktop
| Tool | Voxels | Meshes | Points | Vector fields | Good for |
|---|---|---|---|---|---|
| napari | X | X | X | X | exploring and annotating n-dimensional images; a layer type per data type |
| Fiji + BigDataViewer / BigVolumeBrowser / MoBIE | X | X | X | terabyte microscopy volumes, arbitrary re-slicing, sharing projects | |
| 3D Slicer | X | X | X | X | clinical volumes, DICOM, segmentation, deformation fields |
| ParaView | X | X | X | X | simulation output, glyphs and streamlines, everything at once |
| Blender | OpenVDB only | X | positions only | publication figures and animation; full control of light and material | |
| Blender + Microscopy Nodes | X | X | the same, but it loads OME-TIFF and OME-Zarr stacks directly | ||
| CloudCompare | X | X | arrows only | registering and comparing scans, distances between clouds | |
| MeshLab | X | X | cleaning, repairing and reconstructing surfaces |
the whole list is on the page version of these slides
the whole table is
sitedata/tools.yaml,
and a pull request is the quickest way to add what we have missed.
How to pick a visualization approach
Step 4 - choosing a tool - libraries
| Tool | Voxels | Meshes | Points | Vector fields | Good for |
|---|---|---|---|---|---|
| VTK / PyVista / vedo | X | X | X | X | the general-purpose one; renders headless, so it runs on a cluster |
| napari (as a library) | X | X | X | X | script the viewer you were already clicking in |
| trimesh | occupancy grids | X | X | mesh repair, boolean operations, measuring | |
| Open3D | occupancy grids | X | X | point cloud registration and surface reconstruction | |
| k3d-jupyter, ipyvolume | X | X | X | 3D inside a notebook, next to the code that made it | |
| Dask + Zarr | X | X | not a viewer - what makes data too large for memory openable at all |
the whole list is on the page version of these slides
the whole table is
sitedata/tools.yaml,
and a pull request is the quickest way to add what we have missed.
How to pick a visualization approach
Step 4 - choosing a tool - browser-based software
| Tool | Voxels | Meshes | Points | Vector fields | Good for |
|---|---|---|---|---|---|
| Neuroglancer | X | X | annotations | huge volumes with segmentations; the whole view is a URL | |
| webKnossos | X | X | skeletons | annotating large volumes, collaboratively | |
| Viv / Avivator | X | highly multiplexed microscopy, no server needed | |||
| itk-vtk-viewer | X | X | X | images, meshes and point sets together, via itk-wasm | |
| NiiVue, VolView | X | X | neuroimaging and clinical volumes; easy to embed in a page | ||
| luxar | as Gaussian splats | X | X | lines and tracks | gigabyte volumes and time-lapses from a plain static file host |
| vizarr | X | OME-Zarr straight from a URL - what the IDR serves its images with | |||
| Mol* Volumes & Segmentations | X | X | cryo-EM maps with their segmentations - the viewer behind EMDB and EMPIAR | ||
| Potree | X | streaming point clouds far too large to load | |||
| three.js / Babylon.js | [add-on](https://github.com/Donitzo/three.js-volume-renderer) | X | X | write it yourself | anything custom - the illustrations in these slides are three.js |
the whole list is on the page version of these slides
the whole table is
sitedata/tools.yaml,
and a pull request is the quickest way to add what we have missed.
What this needs from your machine
- The GPU matters more than the CPU. Volume rendering is work per pixel and depends on GPU shaders. Integrated graphics handle small volumes and get slow quickly.
- Graphics memory is the limitation. A volume has to fit in VRAM to be rendered interactively. When it does not, data has to be downsampled, chunked, or streamed.
- In the browser you need WebGL2, which every current browser has. WebGPU is arriving and is considerably faster.
- You can run expensive rendering jobs headless (without a graphical user interface) on the cluster, for example with Blender.