NVIDIA Data Flywheel + Cisco Secure AI Factory
Solution overview
In this lab you will have full access to a deployment of NVIDIA Data Flywheel and learn about the full stack it runs on. The lab is composed of the underlying Cisco Secure AI Factory with NVIDIA infrastructure, the NVIDIA AI Enterprise software stack, and a full stack web application that puts the flywheel to work. We will explain the various layers of the lab allowing you to investigate the role each component brings to the full solution.
Data Flywheel
AI applications in production produce data: prompts, completions, tool calls and user feedback. The NVIDIA data flywheel is a mechanism that allows you to take that data and pivot your AI application to a better, cheaper model while producing the same or better results for your users.
Most enterprises reach for the largest available model when they first put a use case into production, because a large model is the fastest way to get to acceptable accuracy. That decision determines the cost and latency profile of the application for its lifespan. But the needs of many production AI workloads are narrower than the underlying model serving them. A model answering one well-defined question, over and over, on the same kind of prompt, rarely needs 70 billion parameters to do it.
The Data Flywheel addresses this mismatch. It accepts prompt and completion logs from your application, groups them by workload, deduplicates them, and uses class-aware stratified splitting to build balanced evaluation and fine-tuning datasets. It then runs a matrix of experiments across a set of candidate models: the raw model replays production prompts, the same model given few-shot examples drawn from your own traffic, and a LoRA fine-tune of that model. Every result is scored against what the incumbent production model actually returned. What it finally produces is a smaller model with LoRA adapters attached that gives equivalent accuracy.
The results can be dramatic. NVIDIA's own internal testing has identified cases where a data flywheel reduced inference costs by up to 98.6% while maintaining comparable accuracy. In that particular case they found a fine-tuned Llama 3.2 1b could handle a tool-routing task that had been running on Llama 3.1 70b. In this lab, we measured a 1-billion-parameter model (70 times smaller than the comparable foundation model) matching its accuracy on our workload, responding in under half a second, and delivering roughly 15 times the throughput on equivalent hardware.
An important caveat: the flywheel is a flashlight, not an autopilot. It narrows a very large search space down to a handful of credible candidates in a few hours instead of a few weeks. However, to get good results the flywheel must be provided with a representative data set to evaluate against. And in the end deciding what actually goes to production stays a human decision.
Cisco Secure AI Factory with NVIDIA
Running a data flywheel means standing up inference, fine-tuning, evaluation and data generation side by side, sharing a GPU pool, on a network fast enough to keep them fed — and doing it without opening a new attack surface every time a model is retrained.
The Cisco Secure AI Factory with NVIDIA is a modular reference design that addresses all those needs in a single system. It combines Cisco's compute, networking, security and observability with NVIDIA's accelerated computing and AI Enterprise software, plus storage certified by both vendors. WWT operationalizes that design. This lab runs on the shared AI Factory environment in our AI Proving Ground.
The questions this platform layer answers are the ones every IT organization has to answer before an AI project becomes an AI system:
- Where does the training data live, and who can reach it?
- How do fine-tuning jobs and production inference share a finite GPU pool without starving each other?
- What happens to a model's security posture when it is retrained and redeployed automatically?
- How do we observe any of this once it is running?
For this lab, the flywheel runs on Cisco UCS C845A GPU-optimized rack servers in WWT's shared AI Factory environment, which provides a pool of 8x NVIDIA H200, 4x NVIDIA H200 NVL and 4x NVIDIA H100 NVL GPUs. Cisco Nexus networking carries the east-west traffic between the NeMo microservices, and the whole stack is orchestrated on Red Hat OpenShift. OpenShift makes it possible to spin a NIM up for an evaluation run and tear it down again when the run finishes, instead of paying to keep every candidate model resident.
NVIDIA Software Stack
This is the layer that does the work, and each component replaces something a team would otherwise have to build and maintain by hand.
NVIDIA AI Enterprise is the software platform underneath everything. The supported, secured distribution of the NIM and NeMo microservices, the GPU and network operators that make them schedulable on OpenShift, and the enterprise support and CVE remediation that makes this production grade.
NVIDIA Run:ai is used to efficiently allow multiple NIM deployments on single GPUs. Since this lab is on-demand and also multi-instance this allows us to support many running copies on a smaller hardware footprint. Something customer will also want to take advantage of in real-world scenarios to manage inference costs.
NVIDIA NIM is how every model in this lab is served. NIM is a set of easy-to-use microservices for secure, reliable, high-performance AI model inferencing across clouds, data centers and workstations. Each NIM exposes an OpenAI-compatible /v1/chat/completions endpoint, which is what allows the application to treat a 1B model and a 70B model as interchangeable and to hot-swap a freshly trained LoRA adapter underneath a live endpoint without a redeploy.
In this lab you will use:
meta/llama-3.2-1b-instructNIM — the candidate model, served on the H200 pool with a vLLM FP8 LoRA profilemeta/llama-3.2-1b-instructNIM with LoRA adapters — the same base model with the flywheel's fine-tune appliedmeta/llama-3.3-70b-instruct— the incumbent large model, serving as the accuracy baseline
The NeMo microservices platform is what makes the flywheel automatic instead of manual. Programmatic control over datasets, fine-tuning, evaluation and inference means the experiment matrix runs itself:
- NeMo Datastore holds the versioned evaluation and fine-tuning datasets built from logged traffic.
- NeMo Customizer runs the LoRA supervised fine-tuning jobs. This lab uses rank 32, alpha 64 adapters, and a training run over roughly a thousand records completes in about 15 minutes.
- NeMo Evaluator scores every candidate (base, in-context learning and customized) against the incumbent model's actual production responses using an LLM-as-judge similarity metric, and produces a LoRA which the human can decide to deploy.
Around those microservices sits the flywheel service itself: Elasticsearch as the log store where tagged prompt/completion records land, with MongoDB, Redis and Celery handling job metadata and orchestration so a run can be kicked off with a single API call and monitored to completion.
Our implementation follows the NVIDIA AI Blueprint for building data flywheels as its reference architecture.
Full Stack Web Application
The last layer of this lab is the application that gives the flywheel something to optimize. WWT built a clinical coding assistant that reads a discharge summary and predicts the primary and secondary ICD-10 codes. This is a narrow, high-volume, well-defined task, which is precisely the profile of workload where a flywheel pays off.
Our example web application produces ICD-10 medical codes. Clinical records are exactly the kind of data an enterprise cannot casually pipe into a fine-tuning job, so rather than using real patient data we synthetically generated discharge summaries with known ground-truth ICD-10 codes — 36 diagnosis categories spanning cardiac, respiratory, GI, trauma, neurological, overdose and psychiatric presentations. This allowed us to build a statistically useful corpus for a workload without having to handle actual regulated data.
The application runs every case through three models at once (the base 1B, the fine-tuned 1B, and the 70B baseline) and shows you all three responses side by side, validated against the known ground-truth codes and graded as an exact match, a correct ICD-10 category, a clinically related diagnosis, a different diagnostic area, or an invalid code.
You will also drive the flywheel loop yourself. From the application you can generate new synthetic cases, load additional of these generated records into Elasticsearch, kick off a new Data Flywheel training job against everything currently indexed, and then extract the resulting LoRA adapters and hot-swap them into the live NIM endpoint. The end-to-end path is short enough to see in a single sitting:
Synthethic Records → Elasticsearch → Data Flywheel → LoRA adapters → NIM endpoint
Built with a Vue.js front end and a FastAPI backend, the application is an example of modern full stack architecture built for deployment at scale.
At WWT, we understand the need to build solutions that seamlessly fit into our clients' organizational landscape. Some portions of an AI application will always be new, but the other technical decisions don't have to be. Using modern software development practices in tandem with new technologies allows WWT to quickly build solutions that are ready for the enterprise.