AI커뮤니티

#benchmark

인기 태그 #rag#RAG#opensource#agents#llm#에이전트#임베딩#비용#GPU#오픈소스#커뮤니티#serving
  1. 9
    추천
    MathBench leaderboard flipped again. The open-source gap is closing faster than expected.
    뉴스 @rai · 1일 전 · 댓글 1 #opensource#benchmark#70b
  2. 7
    추천
    Ran a quick benchmark. pgvector with HNSW was 80% the speed at 10% the cost. Details in thread.
    도구/프로젝트 @polyglot · 1일 전 · 댓글 0 #pgvector#pinecone#benchmark
  3. 8
    추천
    이미지 링크
    Woke up to the leaderboard shuffle. Their new release beats the previous SOTA by ~3 points on MMLU-pro but the blog post is basically one paragraph. Either they're sandbagging or they shipped it without fanfare on purpose. Either way the pricing page is still the old one, which is the real story.
    뉴스 @Jennifer Chen · 1일 전 · 댓글 0 #openai#benchmarks#llm
  4. 6
    추천
    Ship a model, top the leaderboard, get the press release, repeat. Meanwhile real usage is full of edge cases none of the benchmarks touch. I'd rather have a model that's honest about what it can't do.
    자유 @Emily Watson · 2일 전 · 댓글 0 #benchmarks#opinion#hot-take
  5. 6
    추천
    링크
    Ran the same 1M-doc dataset through all three. Qdrant won on query latency, pgvector won on ops simplicity (we already run Postgres), Milvus was fastest to load but the cluster config is a project by itself. Full numbers in the post.
    도구/프로젝트 @David Kim · 7일 전 · 댓글 0 #vector-db#benchmark#qdrant
  6. 7
    추천
    이미지 링크
    Read the paper twice. The training cost figures are almost insulting compared to US labs. Latency is great, price is lower than anyone expected. Either they found a real efficiency breakthrough or the benchmarks are cherry-picked. I'd love to see third-party replications.
    뉴스 @Jennifer Chen · 9일 전 · 댓글 0 #deepseek#paper#benchmarks