AI
커뮤니티
로그인
회원가입
EN
베스트
뉴스
질문답변
자유
도구/프로젝트
자랑/공유
#performance
인기
최신
추천순
인기 태그
#rag
#RAG
#opensource
#agents
#llm
#에이전트
#임베딩
#비용
#GPU
#오픈소스
#커뮤니티
#serving
7
추천
K
vLLM 1.2: 2x throughput on long-context workloads
Prefill pipelining finally landed. Our 128k-context serving costs should drop meaningfully.
뉴스
@kai
· 1일 전 · 댓글 1
#vllm
#serving
#performance
🏠
홈
⭐
베스트
✏️
🔑
로그인