Một hệ thống LLM gồm những gì

Một hệ thống LLM gồm những gì An architecture diagram generated by Archify. Văn bản · prompt của bạn · Architecture component Văn bản prompt của bạn Tokenizer · vocab + template · Architecture component · sai thì output vô nghĩa Tokenizer vocab + template sai thì output vô nghĩa Architecture · layer, head, hidden · Model artifact — tải cùng một checkpoint · quyết định shape Architecture layer, head, hidden quyết định shape Weights · safetensors · Model artifact — tải cùng một checkpoint · quyết định chất lượng Weights safetensors quyết định chất lượng Task head · classification / LM · Model artifact — tải cùng một checkpoint · quyết định output Task head classification / LM quyết định output Generation · temperature, top-p · Runtime — bạn cấu hình · sai thì lặp hoặc khô Generation temperature, top-p sai thì lặp hoặc khô Memory · KV cache, dtype · Runtime — bạn cấu hình · sai thì hết VRAM Memory KV cache, dtype sai thì hết VRAM Inference engine · batching, streaming · Runtime — bạn cấu hình · quyết định throughput Inference engine batching, streaming quyết định throughput Hardware · CPU / GPU / VRAM · Architecture component · trần cứng Hardware CPU / GPU / VRAM trần cứng chuỗi ký tự ids + mask nạp tham số hidden states logits K, V đã lưu token đã chọn kernel + lịch chạy Model artifact — tải cùng một checkpoint Runtime — bạn cấu hình Legend Frontend Backend Database Cloud Security Message bus External

Nếu output sai

  • • Kiểm tra tokenizer và preprocessing trước tiên
  • • Sau đó tới model head và bước decoding
  • • Đừng đổi model khi lỗi nằm ở lớp biểu diễn

Nếu inference chậm

  • • Nhìn batch size, KV cache và attention kernel
  • • Tách TTFT khỏi tốc độ sinh token để biết nghẽn ở đâu
  • • Kiểm tra memory bandwidth và mức sử dụng GPU

Nếu model không vừa GPU

  • • Hạ precision hoặc dùng model quantized
  • • Offload bớt layer xuống CPU
  • • Hoặc đổi sang inference engine phù hợp hơn