Healthcare · 1 Oct 2026 · 20:31 CEST
Introducing NV-Reason-CT Open 3D CT VLM for Radiologist Chain-of-Thought Reasoning

Publisher preview · OZZZER analysis pending editorial review.
PUBLISHER ARTICLE PREVIEW
From the original article
Radiology AI has made remarkable strides in detecting abnormalities across chest X-rays, pathology slides, and 2D scans. Yet one of the most clinically rich and data-dense modalities—the 3D computed tomography (CT) scan—remains largely underserved by modern vision language models (VLMs). Frontier general-purpose models perform poorly on volumetric imaging, and most open medical AI models lack the multistep conversational depth that radiologists need to trust and verify AI-generated findings.
NVIDIA is addressing this gap with NV-Reason-CT, a VLM purpose-built for 3D CT analysis. NV-Reason-CT extends chain-of-thought reasoning to full volumetric CT, generating structured diagnostic reports, emulating radiologist internal thinking, and supporting multistep follow-up conversation across chest and abdomen. It builds on the reasoning methodology pioneered by NV-Reason-CXR, validated in a multireader clinical study accepted at RSNA 2026 confirming radiologist time savings while maintaining diagnostic accuracy.
NV-Reason-CT is an open research and development foundation; not an autonomous diagnostic system or a cleared clinical product. It is an AI foundation model designed for researchers and developers building specialized CT analysis applications to post-train for their use case.
A single abdominal CT study can comprise 300–600 axial slices, encoding anatomical context across three spatial dimensions that a standard 2D encoder simply cannot reconstruct from independent slices.
This volumetric complexity creates a series of compounding challenges for medical AI to do with perception, reasoning, and conversational depth:
NV-Reason-CT combines a dedicated full 3D vision transformer (ViT) encoder with a language model trained to generate chain-of-thought reasoning that mirrors how radiologists systematically analyze CT volumes.
Unlike approaches that adapt 2D encoders to CT by treating slices independently, NV-Reason-CT processes the CT volume as a true 3D input. This preserves through-plane anatomical continuity and enables the model to reason about structures holistically, the way a radiologist would when scrolling through a study.
The model architecture combines Qwen3.5-4B LLM with 3D ViT (Primus/Colipri). All weights are retrained end-to-end on large cohort or CT data with structured report, reasoning traces, multistep VQA (designed internally). Standard transformer-based VLMs are designed for 2D images. Adapting these to CT by flattening a volume into a sequence of 2D slices loses the spatial structure that defines volumetric pathology.
The encoder architecture is adapted from Primus 3D ViT, initialized with Colipri weights prior to training; It processes CT volumes resampled to 192³
Source
NVIDIA · 1 Oct 2026 · 20:31 CEST
Open the original at NVIDIA ↗