The smaller parameter model sizes of Chinese models (spanning from 200bn to 1.6T parameters, at 2-10% of leading SOTA models, due to constrained access to high-end computing), and highly efficient architectures (MoE, Sparse Attention, OCR etc., at 3-5% activated parameters only vs. total parameter sizes) all contribute to the much lower training and inference costs for Chinese models vs. leading US models.