Một request generation đi qua inference server

Một request generation đi qua inference server A sequence diagram generated by Archify. prompt + sampling params xếp vào hàng đợi prefill cả prompt ghi K, V của prompt logits bước đầu token đầu tiên decode một token K, V cũ dùng lại stream token tiếp request mới chen vào gặp EOS, đóng stream Prefill — đo bằng TTFT Decode — mỗi bước một token Batching liên tục Client · chat API · Sequence participant Client chat API Server · TGI / vLLM · Sequence participant Server TGI / vLLM Scheduler · batching · Sequence participant Scheduler batching GPU · attention · Sequence participant GPU attention KV cache · paged blocks · Sequence participant KV cache paged blocks Sampling · top-p / temp · Sequence participant Sampling top-p / temp Legend request return async trace default message

Hai pha, hai nút thắt khác nhau

  • • Prefill xử lý cả prompt một lượt, nghẽn ở compute
  • • Decode sinh từng token, nghẽn ở memory bandwidth
  • • TTFT đo pha đầu, tokens/s đo pha sau

KV cache và continuous batching

  • • K, V của token cũ không đổi nên được lưu lại
  • • PagedAttention chia cache thành block, giảm phân mảnh
  • • Slot trống được lấp ngay thay vì chờ hết batch