Lab validationThree exact hardware and software configurations
ApexSilica · Accelerator enablement
Production-model qualification across three accelerator platforms
Model and system profiling on Intel Core Ultra, NVIDIA Jetson Orin and Axelera Metis, with retained improvements and rejected experiments recorded against exact configurations.
Production performance depends on how the model graph, compiler, precision, data movement, runtime and serving configuration work together on the target accelerator.
What we delivered
Qualified Qwen3-8B INT4 workloads on Intel Core Ultra 9 285H and NVIDIA Jetson Orin configurations.
Profiled YOLO and RT-DETR-class pipelines on Axelera Metis and documented both successful and blocked paths.
Recorded the exact workload, software, precision and measurement method for every retained result.
What is proven
✓Nearly doubled Intel GPU throughput for a bounded shared-prefix traffic profile and reduced NPU time to first token by 89% for a bounded interactive profile.
✓Reduced Jetson Orin batch-one generation latency by 17% and increased batch-four throughput from approximately 31.5 to 58.9 tokens per second after isolating an FP16 vocabulary projection in the INT4 path.
✓Sustained approximately 565 FPS on the qualified Axelera YOLO11n pipeline and reduced four-stream median latency by 28.6% without reducing throughput.
✓An isolated RMSNorm change was rejected when it slowed the complete model, while RT-DETR profiling identified the next compiler and quantization work required.
Technical methods
Model and graph bring-upEnd-to-end profilingQuantization and compiler diagnosisRuntime and serving optimizationCorrectness and performance qualification
Scope of this evidence
This evidence covers lab validation on the recorded hardware, model, SDK, compiler, runtime, precision, inputs and measurement method. Customer production outcomes are established through a qualified deployment in the target environment.