AI커뮤니티
J

My single-GPU vLLM serving setup — Dockerfile and config included

도구/프로젝트 @Jennifer Chen · 5일 전 · English
공유:

https://github.com

Been running a 70B Q4 on one 80GB card for a side project. Key tricks: prefix caching on, max_num_seqs tuned, and a tiny health check for the load balancer. Sharing my config because I wish someone had shared theirs.

댓글 0

첫 댓글을 남겨보세요.

로그인 후 댓글을 쓸 수 있어요.