Quectel's QSC600XA-AP Puts 24 TOPS and 5 W on an M.2 Card

恒森科技
QuectelM.2 AI acceleratorEdge AIAxera AX8850Embedded computing
Quectel launched the QSC600XA-AP on 29 September 2026, an M.2 2280 B-M Key edge AI board built on Axera's AX8850 SoC: 24 TOPS@INT8 of NPU throughput at 5 W full load, an octa-core Cortex-A55 at 1.5 GHz, 8 GB LPDDR4X standard with a 4 GB option, PCIe 2.0, 22 x 80 x 3.1 mm, roughly 6.37 g bare, and an optional heatsink enclosure. It runs a native Linux 5.15 kernel and is validated on Intel, AMD, Raspberry Pi and Rockchip platforms, with more than 70 open-source models (Qwen3.5, Qwen3-VL, Qwen2.5-VL, YOLO, Whisper) ready to deploy. The value is the socket rather than the silicon: existing gateways, robots, signage players and NVRs gain on-device inference without a motherboard redesign. The board is commercially available for evaluation.

Quectel announced the QSC600XA-AP on 29 September 2026 — an M.2 2280 B-M Key computing board built around Axera's AX8850 SoC. The headline numbers: 24 TOPS@INT8 of NPU throughput, 5 W of NPU power at full load, 8 GB of LPDDR4X standard with a 4 GB option, a PCIe 2.0 host interface, and a board measuring 22 × 80 × 3.1 mm at roughly 6.37 g bare. The form factor matters more than the TOPS figure here: existing gateways, service robots, industrial controllers, digital signage players and NVRs only need a free M.2 slot to gain on-device AI inference, with no motherboard redesign. Quectel says the board has entered commercial availability and can be supplied for evaluation.

What's actually on the board

AX8850 is worth understanding separately from the card. Axera's own product page lists an octa-core Cortex-A55 CPU at 1.5 GHz, an NPU rated at 24 TOPS@INT8, H.264/H.265 codecs capable of 8K@30fps at the chip level, 64-bit LPDDR4/LPDDR4x, eMMC 5.1, SPI Flash, dual Ethernet, one USB 3.0 plus two USB 2.0, dual HDMI outputs, PCIe and SATA — all in a 16 mm × 16 mm package. What Quectel actually ships (Quectel, 29 Sep 2026; Axera product page):

  • NPU: 24 TOPS@INT8 with mixed-precision support; Transformer, ViT and DETR are called out as supported topologies
  • CPU and memory: octa-core Cortex-A55 at 1.5 GHz, 8 GB LPDDR4X standard, 4 GB optional
  • Form factor: M.2 2280 B-M Key, PCIe 2.0, 22 × 80 × 3.1 mm, roughly 6.37 g bare board
  • Power: 5 W for the NPU at full load; a version with a heatsink enclosure is also offered
  • Software: native Linux 5.15 kernel, compatible with Windows, Android and Ubuntu; validated on Intel, AMD, Raspberry Pi and Rockchip platforms

The cost saver is the socket, not the silicon

M.2 2280 is a specification the PC and embedded world already shares — SSDs and wireless modules use the same mechanical and electrical standard. For a customer with an industrial motherboard, an edge gateway or a Raspberry Pi 5, that means the card can be treated as an ordinary PCIe device. What disappears from the schedule is the new motherboard layout, the mechanical and power validation that follows it, and a fresh EMC and certification cycle. On a product already in volume, those line items usually cost far more than the card itself.

Quectel lists edge computing gateways, service robots, industrial control, intelligent security, digital advertising, and — worth noting — routers and home gateways as target applications. That is a reasonable map of products whose main SoC works fine except for AI.

70+ models, and the toolchain behind them

Quectel's pitch is that the card is ready for more than 70 open-source models spanning large language models, multimodal models, image generation and AIGC, naming the Qwen3.5, Qwen3-VL and Qwen2.5-VL families, the YOLO family and Whisper. Customers, in this telling, skip model conversion, quantization tuning and compatibility work entirely.

Axera's public material fills in the other half: its AI processor instruction set covers more than a hundred operators including Conv, Transformer and LSTM, natively targets DeepSeek, Qwen and Llama architectures, and supports mixed precision from 4-bit to 32-bit. The toolchain spans quantization, compilation and deployment and is compatible with PyTorch, TensorFlow and ONNX; Axera claims an order-of-magnitude efficiency gain over conventional GPGPU designs (Axera, 17 Oct 2025).

Demand-side data points the same way. IDC's July 2026 analysis projects 50 million Gen AI PCs and 432 million Gen AI phones shipping worldwide in 2026, and IDC's FutureScape infrastructure forecast expects 80% of enterprises to deploy distributed edge infrastructure by 2027. CIC Insights, in a February 2026 report on edge AI inference chips, puts 2024–2030 CAGRs at 54.9% for robotics, 34.2% for edge vision products and 34.0% for smart automotive.

How it compares: Hailo-8, Orin Nano, and Axera's own card

  • Quectel QSC600XA-AP: AX8850, 24 TOPS@INT8, 5 W NPU, octa-core A55 plus 8 GB LPDDR4X, M.2 2280 B-M Key / PCIe 2.0 (Quectel, 29 Sep 2026)
  • Hailo-8 M.2 module: 26 TOPS INT8, typically 2.5 W, PCIe Gen3 x4, available in M, B+M and A+E keys, extended-temperature version rated -40 to 85 °C; the ET module datasheet separately lists an 8.65 W TDP and 8.25 W maximum at full utilization. It is a pure inference accelerator — no CPU, so the host supplies processor, memory and operating system (Hailo product page and module datasheet)
  • NVIDIA Jetson Orin Nano 8GB: 40 TOPS (INT8), 7–15 W module power, six Cortex-A78AE cores, 8 GB LPDDR5, 69.6 × 45 mm on a 260-pin SO-DIMM connector; the Super variant is rated 67 TOPS (MAXN_SUPER) at 7–25 W (NVIDIA datasheet)
  • Axera's own M.2 card: same AX8850, 2242/2280 form factors, under 8 W by Axera's own figure, expandable from an M.2 card to a PCIe card, validated on Raspberry Pi 5, RK3588 and Intel industrial PCs (Axera)

These three answer different questions. Hailo-8 looks power-efficient largely because it hands the CPU, memory and operating system to the host. Jetson Orin Nano is a complete SoM: 40 TOPS arrives together with a 7–15 W module power budget and a SO-DIMM connector, which makes it a platform change rather than a drop-in card. QSC600XA-AP sits in between, closer to the accelerator end — it carries its own CPU and memory and boots a full operating system while keeping the standard M.2 footprint. The per-watt figures (roughly 4.8 TOPS/W for the Quectel card, 10.4 TOPS/W for Hailo-8 at typical power and 3.15 at maximum, about 2.7 TOPS/W for Orin Nano 8GB at 15 W) are not directly comparable, because the vendors define power differently — the Quectel number is an NPU figure, the Jetson number covers the whole module.

HSY Perspective

Since this card landed, the question we hear most often is the same every time: will it actually work in the box I already have.

Honestly, nobody argues about the 24 TOPS figure. What stops people is more mundane. M.2 comes in B, M, B+M and A+E keys, and the empty slot on an industrial board is often not wired for a compute card — some only carry USB or SATA, with no PCIe lanes routed at all, so the board simply never enumerates. We see that mistake every year, usually discovered after the prototype run, and re-spinning a board is expensive. Thermals and power come next: a 6.37 g bare card looks harmless, but a 5 W NPU sitting next to the CPU, memory and storage on one small board gets warm inside a sealed enclosure. We push customers to think about airflow early, or take the version with the heatsink shell. PCIe 2.0 is also narrower than people assume — if the design feeds several HD streams into the NPU, work out early how decode and scaling are distributed, or you end up with an idle NPU and a saturated bus.

The 70-plus model list matters less than the number suggests. What we care about is that quantization tuning and operator adaptation are already done. Anyone who has deployed at the edge knows that a model "running" and a model "running at acceptable accuracy without operators falling back to the CPU" are two very different claims, and that gap is where schedules usually die. If Axera's toolchain genuinely lands Qwen, YOLO and Whisper straight onto the card, what the customer saves is calendar time, not just development budget.

We carry several routes that customers compare side by side. NVIDIA wins on software ecosystem depth — there is no argument about JetPack or CUDA — but it is an SoM, and carrier redesign and mechanical work are a real project. Hailo's cards have beautiful power numbers, at the cost of needing a host to supply CPU, memory and an operating system. Axera and Quectel take a third path: a complete SoC on a standard card that can act as a standalone node. We usually ask one question first — are you adding an AI co-processor to an existing controller, or replacing that controller with a platform that runs its own OS? Answer that honestly and the part number usually picks itself.

One caveat worth saying out loud: this is not a decision to make from a datasheet. Ask for a board, run it in the actual enclosure at the actual ambient temperature before committing. Send us the M.2 key, the number of routed PCIe lanes and the thermal budget, and we will sanity-check the fit first — that is cheaper than finding out on the production line.

Sources: Quectel, "Quectel's QSC600XA-AP M.2 Smart Computing Board" press release (quectel.com.cn, 29 Sep 2026); Axera AX8850 product page and M.2 accelerator card announcement (axera-tech.com, 17 Oct 2025); Hailo-8 M.2 AI acceleration module product page and Hailo-8 M.2 Key M ET module datasheet Rev2.2 (hailo.ai); NVIDIA Jetson Orin Nano series datasheet and product page (nvidia.com); IDC, edge AI analysis and FutureScape infrastructure predictions (idc.com, 15 Jul 2026); CIC Insights, global and China edge AI inference chip market forecast 2024–2030 (cninsights.com, 13 Feb 2026).