AI커뮤니티

베스트

  1. 15
    추천
    폐쇄망용 완전 오프라인 RAG. Ollama만으로 동작, 임베딩은 bge-m3. 피드백 환영합니다.
    도구/프로젝트 @sean · 1일 전 · 댓글 3 #RAG#폐쇄망#Ollama
  2. 14
    추천
    이미지 링크
    4bit 양자화가 기본 지원돼서 이제 노트북에서도 돌아간다고 하네요. 27B가 16GB VRAM에서 무리 없이 돌아가는 수준이라니. 로컬 AI가 현실로.
    뉴스 @nova · 4시간 전 · 댓글 2 #Gemma#오픈소스#구글
  3. 13
    추천
    이미지 링크
    도메인 지식이 필요하면 RAG, 스타일/포맷 고정이면 FT라는 기본 공식이 있는데, 실제 프로젝트에서는 어떤 기준으로 결정하시나요?
    질문답변 @polyglot · 13시간 전 · 댓글 0 #RAG#파인튜닝#아키텍처
  4. 12
    추천
    가격 대비 성능이 미쳤다. $0.14/M in. 로컬 vLLM 대신 API로 갈아탈까 고민 중.
    뉴스 @gptkim · 1일 전 · 댓글 3 #DeepSeek#벤치마크#가격
  5. 12
    추천
    링크
    2.8T 파라미터면 DeepSeek V4 Pro보다 75% 크다고 합니다. 1M 컨텍스트 + 항상 켜진 추론 모드. 오픈소스가 이제 안 따라잡히네요.
    뉴스 @mlpark · 21시간 전 · 댓글 2 #KimiK3#오픈소스#Moonshot
  6. 12
    추천
    이미지 링크
    8B~27B 모델을 서빙하려고 합니다. throughput이 중요하면 vLLM, VRAM이 빠듯하면 llama.cpp라는 얘기가 있는데 실제 운영하시는 분들 경험 공유 부탁드려요.
    질문답변 @stronguser · 9시간 전 · 댓글 1 #서빙#vLLM#llama.cpp
  7. 12
    추천
    이미지 링크
    사내 매뉴얼 2천 페이지를 RAG로 묶어서 챗봇을 만들었습니다. 임베딩은 bge-m3, 서빙은 vLLM으로 했고, 정확도는 사내 QA 기준 87% 나옵니다.
    자랑/공유 @hanbit · 6시간 전 · 댓글 1 #RAG#사내#배포
  8. 11
    추천
    L4 한 장으로 Qwen2.5-7B 서빙. 처리량 만족스럽습니다. 양자화 없이도 충분하네요.
    자랑/공유 @mlpark · 1일 전 · 댓글 2 #vLLM#서빙#GPU
  9. 11
    추천
    이미지 링크
    드디어 에이전트 전용 API가 나왔네요. MCP 네이티브라서 도구 연결이 표준화될 듯. 기존 함수 콜링보다 훨씬 간결해 보입니다.
    뉴스 @gptkim · 5시간 전 · 댓글 1 #OpenAI#에이전트#MCP
  10. 11
    추천
    링크
    코딩 테스트 통과율이 예전보다 확실히 높아진 느낌. AI 도구 사용을 허용할지, 아니면 역량 평가 방식을 바꿔야 할지 고민입니다.
    자유 @nova · 16시간 전 · 댓글 0 #면접#채용#AI도구
  11. 11
    추천
    이미지 링크
    지식베이스 검색 시간이 하루 2시간 → 10분으로 줄었습니다. 초기엔 환각이 문제였는데, 인용 소스 표시를 강제하니 신뢰도가 확 올라갔어요.
    자랑/공유 @alexus · 17시간 전 · 댓글 1 #KB#RAG#생산성
  12. 10
    추천
    LangGraph, CrewAI, 자체구현... 다들 뭐 쓰시는지 궁금합니다.
    자유 @hanbit · 1일 전 · 댓글 3 #에이전트#프레임워크
  13. 10
    추천
    이미지 링크
    Tool-use 지연이 300ms까지 줄었다고 하네요. 에이전트가 실제 업무에 쓰이려면 이 정도 응답 속도가 필요했죠. 기대됩니다.
    뉴스 @mlpark · 8시간 전 · 댓글 0 #Anthropic#에이전트#성능
  14. 10
    추천
    링크
    LangGraph 같은 프레임워크 쓰다가 직접 while 루프로 짜는 게 낫다는 글을 봤는데요. 디버깅이 쉽고 의존성이 적다는 게 이유라네요. 다들 어떻게 생각하세요?
    자유 @sean · 5시간 전 · 댓글 1 #에이전트#프레임워크#의견
  15. 10
    추천
    이미지 링크
    로컬 RAG 서비스 최소 구성 템플릿을 만들어봤습니다. docker compose 하나로 Ollama + FastAPI + ChromaDB가 뜹니다. 피드백 환영해요.
    도구/프로젝트 @sean · 11시간 전 · 댓글 1 #RAG#템플릿#Docker
  16. 9
    추천
    Ollama on a laptop is great for tinkering, but once you need reliability, hosted APIs win on latency and uptime. My experience so far.
    뉴스 @alexus · 1일 전 · 댓글 2 #llm#api#ollama
  17. 9
    추천
    DeepSeek 캐시 히트율 66% 나오는데, 프롬프트 구조 어떻게 짜면 캐시가 잘 맞나요?
    질문답변 @lobster · 1일 전 · 댓글 2 #프롬프트#캐시#비용
  18. 9
    추천
    MathBench leaderboard flipped again. The open-source gap is closing faster than expected.
    뉴스 @rai · 1일 전 · 댓글 1 #opensource#benchmark#70b
  19. 9
    추천
    웹에서 문서를 읽어와서 요약하는 기능인데, 문서에 악성 지시문이 있으면 대응이 어렵네요. 방어 전략 공유 부탁드려요.
    질문답변 @lobster · 1일 전 · 댓글 1 #보안#프롬프트인젝션#방어
  20. 9
    추천
    보안 때문에 전부 온프레미스로. bge-m3 + Qwen2.5-14B로 구축했고, 사용자 만족도 4.2/5 나왔어요. 후기 공유합니다.
    자랑/공유 @mlpark · 22시간 전 · 댓글 1 #사내#챗봇#온프레미스
  21. 9
    추천
    Ticket deflection went from 0% to 38% in a month. Here's the architecture and the mistakes we made.
    자랑/공유 @alexus · 1일 전 · 댓글 1 #support#rag#case-study
  22. 9
    추천
    링크
    솔직히 이번 건 좀 궁금하네요. 정부가 어떻게 안전하다고 결론 내렸는지 과정이 투명하지 않다는 지적. 원문 공유합니다.
    뉴스 @gptkim · 20시간 전 · 댓글 1 #OpenAI#정부규제#프론티어
  23. 9
    추천
    이미지 링크
    국내 LLM 인프라 스타트업이 시리즈 B로 500억을 유치했다고 합니다. 한국어 특화 모델 수요가 확실히 늘고 있는 것 같아요.
    뉴스 @hanbit · 7시간 전 · 댓글 0 #투자#한국#스타트업
  24. 9
    추천
    이미지 링크
    bge-m3에서 최신 모델로 바꾸려고 하는데 벡터 차원이 달라서 재임베딩이 필요할 것 같습니다. 100만 건 정도인데 시간이 얼마나 걸릴지, 꿀팁 있으면 알려주세요.
    질문답변 @lobster · 11시간 전 · 댓글 0 #임베딩#벡터DB#마이그레이션
  25. 9
    추천
    이미지 링크
    벡터 DB 오픈소스 ChromaDB가 1.0을 발표했습니다. 필터링 성능 개선과 인덱스 안정성이 핵심이라고 하네요. 기존 0.4에서 마이그레이션 해보신 분 계신가요?
    도구/프로젝트 @mlpark · 4시간 전 · 댓글 0 #ChromaDB#벡터DB#오픈소스
  26. 9
    추천
    이미지 링크
    인터넷이 안 되는 환경에서 LLM+RAG를 돌려야 하는 프로젝트였는데, Ollama + 로컬 임베딩으로 구성했습니다. 오프라인 배포에 관심 있으신 분들께 도움이 되길.
    자랑/공유 @sean · 8시간 전 · 댓글 1 #오프라인#RAG#사례
  27. 9
    추천
    이미지 링크
    They released the evals and methodology this time, which is rare. HumanEval pass@1 is a couple points ahead of Llama 4.5 and it's faster to serve. We're spinning up a test pod this week. If the license is as permissive as they say, this is a real option for EU teams.
    뉴스 @Emily Watson · 4일 전 · 댓글 0 #mistral#moe#code
  28. 8
    추천
    한국어 문서 검색인데 bge-m3 vs multilingual-e5 고민입니다. 경험 공유 부탁해요.
    질문답변 @vectorlee · 1일 전 · 댓글 4 #RAG#임베딩
  29. 8
    추천
    로컬에서 돌릴 수 있는 새 오픈웨이트 모델 소식. 7B인데 성능이 무섭다.
    뉴스 @kai · 2일 전 · 댓글 1 #LLM#오픈소스
  30. 8
    추천
    드디어 에이전트 전용 API가 나왔네요. MCP 서버가 기본 지원이라 도구 연동이 쉬워질 듯. 개발자 문서부터 읽어봐야겠어요.
    뉴스 @gptkim · 22시간 전 · 댓글 2 #OpenAI#에이전트#MCP
  31. 8
    추천
    도메인 용어가 많은데 파인튜닝이 나을지 RAG가 나을지. 둘 다 해보신 분들 의견 궁금합니다.
    질문답변 @gptkim · 1일 전 · 댓글 1 #파인튜닝#RAG#선택
  32. 8
    추천
    Every framework becomes a debugging nightmare past 500 lines. Plain code + function calling has been better for us.
    자유 @polyglot · 1일 전 · 댓글 1 #agents#frameworks#opinion
  33. 8
    추천
    비용/지연/프롬프트 버전까지 한 번에 보이니까 디버깅 시간이 반으로 줄었습니다. 무료 티어로 충분해요.
    도구/프로젝트 @gptkim · 1일 전 · 댓글 2 #Langfuse#트레이싱#모니터링
  34. 8
    추천
    듀얼 3090으로 70B 돌리려다 이틀 삽질. 결국은 OLLAMA_NUM_GPU + 스플릿 모드 조합으로 해결. 정리해봤습니다.
    도구/프로젝트 @nova · 1일 전 · 댓글 0 #Ollama#멀티GPU#70B
  35. 8
    추천
    금융권 고객사에 설치했습니다. 인터넷 완전 차단 환경에서도 임베딩+검색+생성 전부 동작. 자세한 내용은 글에.
    자랑/공유 @sean · 1일 전 · 댓글 1 #폐쇄망#RAG#배포
  36. 8
    추천
    Replaced a custom fine-tune with a generic 8B + good retrieval. Cheaper to maintain, better scores. Data point for you.
    자랑/공유 @vectorlee · 1일 전 · 댓글 1 #rag#finetune#case-study
  37. 8
    추천
    링크
    Mira Murati가 세운 회사가 시드로 2조원대 유치. a16z 리드. 시장 분위기가 확실히 달라졌네요.
    뉴스 @hanbit · 22시간 전 · 댓글 1 #투자#MiraMurati#시드
  38. 8
    추천
    이미지 링크
    코드 문서는 300토큰, 일반 문서는 500~800토큰으로 쓰고 있는데요. 법률 문서처럼 구조화된 건 섹션 단위로 자르는 게 낫다는 의견도 있더라고요. 다들 어떻게 하세요?
    질문답변 @rai · 6시간 전 · 댓글 1 #RAG#청크#임베딩
  39. 8
    추천
    링크
    사내용 챗봇 하나 돌리는데 GPU 서버 빌리는 게 월 200만원이 넘네요. 트래픽이 적으면 API 쓰는 게 이득이라는 의견도 있고요. 실제 사례 있으면 공유 부탁드려요.
    자유 @stronguser · 20시간 전 · 댓글 0 #비용#인프라#GPU
  40. 8
    추천
    이미지 링크
    프롬프트/비용/지연을 한 대시보드에서 보니 운영 난이도가 확실히 내려갔습니다. 셀프호스팅도 가능해서 사내 도입에 부담이 없네요.
    도구/프로젝트 @alexus · 9시간 전 · 댓글 0 #Langfuse#옵저버빌리티#추적
  41. 8
    추천
    이미지 링크
    매일 아침 AI 관련 뉴스 5개를 요약해서 텔레그램으로 보내주는 봇을 만들었습니다. DeepSeek로 요약하고 키워드로 분류하니 꽤 유용하네요.
    자랑/공유 @nova · 10시간 전 · 댓글 0 #텔레그램#뉴스#봇
  42. 8
    추천
    이미지 링크
    Woke up to the leaderboard shuffle. Their new release beats the previous SOTA by ~3 points on MMLU-pro but the blog post is basically one paragraph. Either they're sandbagging or they shipped it without fanfare on purpose. Either way the pricing page is still the old one, which is the real story.
    뉴스 @Jennifer Chen · 1일 전 · 댓글 0 #openai#benchmarks#llm
  43. 8
    추천
    이미지 링크
    Weights are up on the hub. The 8B with the new tokenizer is suspiciously good at tool calling for its size. Fine-tuning on our internal data took 6 hours on a single H100. If you need a cheap default for agents, this is hard to beat.
    뉴스 @Tom Hall · 7일 전 · 댓글 0 #meta#llama#opensource
  44. 8
    추천
    Our org has 10k pages of policy PDFs nobody reads. Built a RAG bot with citations that links back to the exact page. Support queries dropped noticeably and compliance is happy because everything is traceable. Side project that became the team's favorite tool.
    자랑/공유 @Jennifer Chen · 2일 전 · 댓글 0 #rag#chatbot#showcase
  45. 7
    추천
    Feed it a CSV export, it finds anomalies and writes a Korean explanation via an LLM. Useful for MSP folks.
    도구/프로젝트 @sean · 1일 전 · 댓글 1 #cloud#cost#cli
  46. 7
    추천
    가성비 미쳤습니다. 4bit로 12GB 램 노트북에서 돌아간다고 하니 로컬 실험용으로 딱이네요.
    뉴스 @mlpark · 1일 전 · 댓글 1 #Gemma#양자화#로컬LLM
  47. 7
    추천
    Prefill pipelining finally landed. Our 128k-context serving costs should drop meaningfully.
    뉴스 @kai · 1일 전 · 댓글 1 #vllm#serving#performance
  48. 7
    추천
    현재 Chroma 쓰다가 Qdrant로 옮길지 고민 중. 데이터 500만 건 정도인데 경험 있으신 분?
    질문답변 @vectorlee · 1일 전 · 댓글 1 #vectordb#마이그레이션#Qdrant
  49. 7
    추천
    bge-m3에서 다른 모델로 바꾸려는데, 기존 문서 전부 다시 인덱싱해야 하는 게 맞는지요?
    질문답변 @mlpark · 1일 전 · 댓글 1 #임베딩#인덱싱#RAG
  50. 7
    추천
    Mine: Claude for code, DeepSeek for bulk tasks, local Ollama for privacy stuff. What's yours?
    자유 @alexus · 1일 전 · 댓글 1 #stack#daily#workflow
  51. 7
    추천
    일 10만 요청 기준 DeepSeek API vs 자체 GPU. 대략 어디서 크로스오버되나요? 계산해보신 분?
    자유 @mlpark · 1일 전 · 댓글 1 #비용#GPU#API
  52. 7
    추천
    네임스페이스 API가 안정화됐고 성능도 올라갔다고 합니다. 마이그레이션 가이드 공유합니다.
    도구/프로젝트 @vectorlee · 23시간 전 · 댓글 1 #ChromaDB#벡터DB#릴리즈
  53. 7
    추천
    Ran a quick benchmark. pgvector with HNSW was 80% the speed at 10% the cost. Details in thread.
    도구/프로젝트 @polyglot · 1일 전 · 댓글 0 #pgvector#pinecone#benchmark
  54. 7
    추천
    All data stays on device. Weekly digest with topics, mood, and action items. 400 users in 2 weeks!
    자랑/공유 @kai · 1일 전 · 댓글 1 #local-first#journal#product
  55. 7
    추천
    동일 7B 모델로 비교했는데 p50은 비슷하고 p99는 TRT-LLM이 더 안정적이었습니다. 수치 표로 정리했어요.
    자랑/공유 @gptkim · 1일 전 · 댓글 1 #vLLM#TensorRT#벤치마크
  56. 7
    추천
    링크
    Multiple model sizes, powering products across Meta. The open-vs-closed debate keeps heating up.
    뉴스 @rai · 1일 전 · 댓글 1 #Meta#Llama3#opensource
  57. 7
    추천
    링크
    Updated list of open LLMs you can use commercially. Handy when picking a model for a client project.
    도구/프로젝트 @sean · 21시간 전 · 댓글 1 #opensource#llm#list
  58. 7
    추천
    이미지 링크
    시행 시점이 유예되고 중소기업 부담이 줄었다고 합니다. 유럽 스타트업들 숨통 트이는 분위기. 국내 규제 논의에도 참고할 만하네요.
    뉴스 @vectorlee · 10시간 전 · 댓글 0 #규제#EU#정책
  59. 7
    추천
    이미지 링크
    프롬프트 캐시 쓰면 50%는 아낀다고 들었는데, 실제로 얼마나 효과 보셨나요? 컨텍스트 압축이나 요약 전략도 궁금합니다.
    질문답변 @kai · 12시간 전 · 댓글 0 #컨텍스트#캐시#비용
  60. 7
    추천
    이미지 링크
    둘 다 써봤는데, Open-webui는 Ollama 연동이 편하고 LibreChat는 멀티모델 관리가 좋았습니다. RAG 기능은 LibreChat가 좀 더 성숙한 느낌? 여러분 선택은?
    도구/프로젝트 @rai · 7시간 전 · 댓글 1 #UI#셀프호스팅#비교
  61. 7
    추천
    이미지 링크
    매주 LLM 논문 하나씩 리뷰하는 블로그를 6개월째 운영 중입니다. 방문자가 꾸준히 늘고 있고, 이메일 구독도 300명 넘었네요. 꾸준함이 답이긴 하네요.
    자랑/공유 @polyglot · 12시간 전 · 댓글 1 #블로그#논문#기록
  62. 7
    추천
    이미지 링크
    Downloaded the 27B variant last night. Quantized to Q4 it fits in 16GB RAM and generates at a very usable speed. Not Claude-level reasoning, but for a local model it's a big step. Google is winning the 'open weights you can actually run' game right now.
    뉴스 @Marcus Obediah · 2일 전 · 댓글 0 #google#open-weights#local-llm
  63. 7
    추천
    이미지 링크
    Read the paper twice. The training cost figures are almost insulting compared to US labs. Latency is great, price is lower than anyone expected. Either they found a real efficiency breakthrough or the benchmarks are cherry-picked. I'd love to see third-party replications.
    뉴스 @Jennifer Chen · 9일 전 · 댓글 0 #deepseek#paper#benchmarks
  64. 7
    추천
    Day 1: I felt like a god. Day 12: I found a bug the AI introduced that took 3 hours to hunt. Day 30: net positive, probably. The tooling is incredible but you still need to understand what you're building. No replacement, just acceleration.
    자유 @Anna Nowak · 5일 전 · 댓글 0 #solo-dev#ai-tools#review
  65. 7
    추천
    링크
    50 lines of TypeScript, points at our search API, and now everyone queries docs from their editor. It's dumb but it changed how we work. The protocol is clunky but the integration win is real.
    도구/프로젝트 @Marcus Obediah · 4일 전 · 댓글 0 #mcp#internal-tools#docs
  66. 7
    추천
    Collected 3 years of my blog posts and newsletter, fine-tuned a 7B with LoRA. The style imitation is uncanny — my wife thought I wrote the drafts. Uses: first drafts of posts I'm too lazy to start. It's my ghostwriter now.
    자랑/공유 @Marcus Obediah · 4일 전 · 댓글 0 #finetuning#writing#lora
  67. 7
    추천
    12B model on the family PC, a simple web UI. It explains math problems step by step instead of just giving answers. No accounts, no data leaving the house, and no monthly subscription. The kids actually use it.
    자랑/공유 @Raj Patel · 8일 전 · 댓글 0 #local-llm#education#family
  68. 6
    추천
    Chroma vs Qdrant vs LanceDB for a small self-hosted RAG service. Disk and RAM matter — running on a free-tier VM.
    질문답변 @kai · 1일 전 · 댓글 2 #vectordb#rag
  69. 6
    추천
    Scrapes arXiv + HN, summarizes with an LLM, posts as threads. 200 stars in a week. AMA.
    자랑/공유 @rai · 1일 전 · 댓글 2 #newsletter#automation
  70. 6
    추천
    국내+글로벌 AI 논의를 한 곳에서. 짧게 쓰레드처럼 쓰는 곳입니다. 자유롭게 이야기해요.
    자유 @nova · 2일 전 · 댓글 2 #인사#커뮤니티
  71. 6
    추천
    Big claim. If real, this changes agent UX a lot. Anyone benchmarked it yet?
    뉴스 @alexus · 1일 전 · 댓글 1 #anthropic#latency#agents
  72. 6
    추천
    오프라인 어시스턴트가 실생활에 들어오는 시점이네요. 개인정보 처리 관점에서도 의미 있는 변화.
    뉴스 @nova · 1일 전 · 댓글 1 #온디바이스#AI폰#로컬
  73. 6
    추천
    LLM이 스트리밍으로 출력할 때 JSON을 점진적으로 파싱하는 방법 질문입니다. partial JSON 파서 쓰시는 분 계신가요?
    질문답변 @hanbit · 23시간 전 · 댓글 2 #streaming#json#파싱
  74. 6
    추천
    Long agent sessions blow up tokens. Do you use summarization, sliding windows, or something else?
    질문답변 @kai · 1일 전 · 댓글 1 #context#agents#tokens
  75. 6
    추천
    Model cites chunks that don't actually support the answer. Any post-processing tricks that work?
    질문답변 @rai · 1일 전 · 댓글 0 #rag#hallucination#citations
  76. 6
    추천
    LLM을 잘 쓰는 것만으로는 부족하다는 말이 많아지네요. 시스템 설계? 벡터 DB? MLOps? 다들 뭐가 우선이라 생각하세요?
    자유 @hanbit · 1일 전 · 댓글 1 #커리어#공부#개발자
  77. 6
    추천
    보안팀이랑 임원분들이 '아직 이르다'는 입장인데, 실무는 이미 AI 없이는 일이 안 됩니다. 설득 팁 있을까요?
    자유 @lobster · 1일 전 · 댓글 0 #회사#도입#보안
  78. 6
    추천
    With better rerankers and hybrid search, raw embedding quality matters less every quarter. Agree?
    자유 @sean · 1일 전 · 댓글 0 #embeddings#trend#hot-take
  79. 6
    추천
    We went LibreChat for multi-user + RBAC. Open-webui has nicer UX though. Your pick?
    도구/프로젝트 @kai · 1일 전 · 댓글 1 #self-host#chat-ui#open-source
  80. 6
    추천
    드래그앤드롭으로 LLM + 검색 + 슬랙 연동까지. 개발자도 가끔은 로우코드가 편하네요.
    도구/프로젝트 @hanbit · 1일 전 · 댓글 1 #n8n#자동화#워크플로
  81. 6
    추천
    CSV of prompts → JSONL of responses + scores. 100 lines of Go. Maybe useful for your eval pipeline.
    도구/프로젝트 @alexus · 1일 전 · 댓글 0 #cli#evals#opensource
  82. 6
    추천
    Scrapes 20 sources, dedupes, summarizes, posts at 8am. 1.2k subscribers now. AMA about the pipeline.
    자랑/공유 @rai · 1일 전 · 댓글 1 #telegram#bot#newsletter
  83. 6
    추천
    주 1회 LLM 논문 리뷰. 누적 12만 뷰 달성. 글쓰기 + 실험 재현을 병행하니 배우는 게 두 배네요.
    자랑/공유 @hanbit · 1일 전 · 댓글 0 #블로그#논문#리뷰
  84. 6
    추천
    링크
    Doctors' AI scribe just raised a huge round at $5.3B valuation. Healthcare AI is the real money-maker apparently.
    뉴스 @alexus · 23시간 전 · 댓글 0 #healthcare#ai#funding
  85. 6
    추천
    링크
    뉴스 매체 AI 생성 콘텐츠 비율 추적 연구입니다. 지역/대학 언론이 10배 가까이 증가했다고 하네요.
    뉴스 @nova · 1일 전 · 댓글 1 #논문#뉴스#GenAI
  86. 6
    추천
    링크
    개발자용 월간 AI 뉴스 큐레이션 레포입니다. 매월 최신 LLM/도구 정리돼 있어서 스크랩해두면 좋아요.
    도구/프로젝트 @vectorlee · 23시간 전 · 댓글 1 #뉴스#큐레이션#GitHub
  87. 6
    추천
    링크
    보안 우려로 LLM 도입을 꺼리시는 분들이 있는데, 폐쇄망 배포 사례를 보여드리면 좀 나아질까요? 비슷한 경험 있으신 분들의 조언 부탁드립니다.
    자유 @hanbit · 14시간 전 · 댓글 0 #도입#조직#설득
  88. 6
    추천
    이미지 링크
    이메일 요약 → 슬랙 전송 → 주간 리포트 생성까지 n8n으로 묶어봤습니다. 노코드인데 LLM 노드가 잘 돼 있어서 생각보다 강력하네요.
    도구/프로젝트 @kai · 15시간 전 · 댓글 0 #n8n#자동화#워크플로
  89. 6
    추천
    이미지 링크
    Just got the email. Cache writes are 50% cheaper and cache reads dropped too. For anyone running long agent loops this is basically free money. We're re-running our cost model tonight but I suspect this changes the math on long-context agents.
    뉴스 @Sofia Alvarez · 3일 전 · 댓글 0 #anthropic#pricing#agents
  90. 6
    추천
    이미지 링크
    Ordered 8 cards for our inference cluster two weeks ago. Sales rep just told me new orders are looking at Q4. Everyone is doing the same math: cheap tokens need efficient hardware, and there isn't enough to go around. If you're planning capacity for next year, order now.
    뉴스 @Raj Patel · 6일 전 · 댓글 0 #nvidia#gpu#infra
  91. 6
    추천
    Ship a model, top the leaderboard, get the press release, repeat. Meanwhile real usage is full of edge cases none of the benchmarks touch. I'd rather have a model that's honest about what it can't do.
    자유 @Emily Watson · 2일 전 · 댓글 0 #benchmarks#opinion#hot-take
  92. 6
    추천
    Two months in, no regrets so far. Working on an open agent framework and living off savings. The open source ecosystem is chaotic but the pace is addictive. Happy to answer questions about the leap, money, or what I actually do all day.
    자유 @Emily Watson · 4일 전 · 댓글 0 #career#opensource#ama
  93. 6
    추천
    링크
    We were merging PRs blind. Now every RAG change runs 40 questions with expected answers and fails the build if recall drops. It's crude but it caught two regressions last week. Sometimes the simple stuff wins.
    도구/프로젝트 @Anna Nowak · 2일 전 · 댓글 0 #rag#evals#testing
  94. 6
    추천
    링크
    Ran the same 1M-doc dataset through all three. Qdrant won on query latency, pgvector won on ops simplicity (we already run Postgres), Milvus was fastest to load but the cluster config is a project by itself. Full numbers in the post.
    도구/프로젝트 @David Kim · 7일 전 · 댓글 0 #vector-db#benchmark#qdrant
  95. 6
    추천
    Workshop is loud, dusty, and has no wifi worth trusting. Pi 5 + a local model + a cheap mic array. It answers questions about my tools, sets timers, and reads specs out loud. Everything runs on-device and it cost under $200.
    자랑/공유 @Sofia Alvarez · 5일 전 · 댓글 0 #voice#offline#hardware
  96. 6
    추천
    링크
    Been dumping notes, highlights, and bookmarks into a local RAG setup for two years. Search is instant and surprisingly good. Open-sourced the whole thing so it's reproducible. Warning: your old notes are more embarrassing than you remember.
    자랑/공유 @David Kim · 6일 전 · 댓글 0 #second-brain#rag#opensource
  97. 6
    추천
    Scrapes my git commits and PRs, summarizes them into 3 bullet points, emails them before standup. Takes 30 seconds of my morning. My manager thinks I'm extremely organized. I'm just lazy with extra steps.
    자랑/공유 @Anna Nowak · 9일 전 · 댓글 0 #automation#agents#humor
  98. 6
    추천
    It joins the call with a normal name, takes notes, and posts them to our wiki. People forget it's there until the notes appear. Uses a tiny local model for transcription. The 'is this legal' question came up once. HR said fine, so we're shipping it.
    자랑/공유 @Marcus Obediah · 1일 전 · 댓글 0 #meetings#notes#side-project
  99. 5
    추천
    I keep coming back to plain code + function calling. Frameworks add too much magic for debugging.
    자유 @polyglot · 1일 전 · 댓글 1 #agents#function-calling
  100. 5
    추천
    RAG 기반 기업 검색 솔루션으로 유치했다고 합니다. 국내도 이제 본격적으로 크네요.
    뉴스 @hanbit · 1일 전 · 댓글 1 #AI#스타트업#투자
  101. 5
    추천
    Small enough to self-host, strong enough for real agent loops. Worth a look for cost-sensitive teams.
    뉴스 @polyglot · 1일 전 · 댓글 0 #mistral#function-calling#self-host
  102. 5
    추천
    We use RAGAS but the scores don't correlate with human judgment. What do you actually use in production?
    질문답변 @polyglot · 1일 전 · 댓글 1 #rag#evaluation#ragas
  103. 5
    추천
    vLLM vs TGI vs llama.cpp server. For a small team serving one 8B model, what's the least ops burden?
    질문답변 @alexus · 1일 전 · 댓글 1 #self-host#serving#llm
  104. 5
    추천
    임베딩 + 소형 LLM만으로 API 없이 운영 가능한지 궁금합니다. 트래픽은 하루 1천 건 수준이에요.
    질문답변 @sean · 1일 전 · 댓글 0 #GPU#임베딩#비용
  105. 5
    추천
    코딩 문제는 다들 잘 풀어요. 문제는 '왜'를 물었을 때. 면접 방식 자체를 바꿔야 하나 고민입니다.
    자유 @nova · 1일 전 · 댓글 0 #면접#신입#에세이
  106. 5
    추천
    I used to read every arxiv paper. Now I wait for good summaries. Am I falling behind?
    자유 @rai · 1일 전 · 댓글 0 #papers#learning#arxiv
  107. 5
    추천
    새 모델 나올 때마다 다같이 벤치마크 까보는 게 제일 재밌네요. 여기 오시는 분들은 뭐 보러 오세요?
    자유 @gptkim · 1일 전 · 댓글 0 #커뮤니티#잡담
  108. 5
    추천
    I packaged my go-to RAG setup into a template. Docker compose up and you have a working API in 2 minutes.
    도구/프로젝트 @sean · 1일 전 · 댓글 1 #docker#ollama#fastapi#template
  109. 5
    추천
    지금은 Git으로 관리하는데, 버전별 A/B 테스트가 필요해졌습니다. 전용 도구 추천 부탁해요.
    도구/프로젝트 @mlpark · 1일 전 · 댓글 1 #프롬프트#관리#도구
  110. 5
    추천
    Gradio got multimodal chat components. Streamlit is catching up. For internal demos, which wins?
    도구/프로젝트 @rai · 1일 전 · 댓글 1 #gradio#streamlit#demo
  111. 5
    추천
    Visual node editor for agent pipelines with live tracing. MIT licensed, would love feedback.
    자랑/공유 @polyglot · 1일 전 · 댓글 0 #opensource#agents#playground
  112. 5
    추천
    매주 토요일 2시간. 30명에서 시작해 지금 12명의 코어 멤버로. 꾸준함의 가치를 배웠습니다.
    자랑/공유 @nova · 1일 전 · 댓글 1 #스터디#모임#커뮤니티
  113. 5
    추천
    링크
    34 of the 100 fastest-growing companies are AI. The capital shift is real and it's accelerating.
    뉴스 @polyglot · 1일 전 · 댓글 1 #funding#report#trend
  114. 5
    추천
    링크
    Both raised big rounds in 2025. If you're building on top of LLMs, infra plays look safer than apps.
    뉴스 @kai · 1일 전 · 댓글 0 #startups#funding#infra
  115. 5
    추천
    링크
    Keep this tab open. Everything from o1 to GPT-5-Codex-Mini lands here first.
    도구/프로젝트 @alexus · 1일 전 · 댓글 1 #openai#changelog#models
  116. 5
    추천
    링크
    여기처럼 짧은 글 위주로 활발한 곳이 또 있을까요? 디스코드/슬랙 커뮤니티도 추천해주시면 감사하겠습니다.
    자유 @vectorlee · 18시간 전 · 댓글 1 #커뮤니티#모임#추천
  117. 5
    추천
    이미지 링크
    The grace period is ending. If you ship anything that touches biometrics or critical infrastructure you need to have your paperwork sorted. Small SaaS teams are mostly fine but the compliance checklist is long and boring. Start with the Annex III list, don't assume you're exempt.
    뉴스 @David Kim · 5일 전 · 댓글 0 #eu-ai-act#regulation#compliance
  118. 5
    추천
    이미지 링크
    The 3B model handles summarization and smart replies locally, and it's snappy. Privacy story is nice. Devs can call it via a new CoreML API. It's not going to replace cloud models for hard tasks, but for simple stuff it's genuinely useful and free.
    뉴스 @Marcus Obediah · 10일 전 · 댓글 0 #apple#on-device#ios
  119. 5
    추천
    Building evals feels like a second product. We tried LLM-as-judge with a strong model and it's decent but biased toward long answers. Anyone using lightweight checks + sampling? What's your minimal viable eval?
    질문답변 @Raj Patel · 4일 전 · 댓글 0 #evals#llm-judge#testing
  120. 5
    추천
    We're about to onboard our first enterprise customers and I have to decide. Shared index with metadata filters seems simpler but I'm scared of cross-tenant leakage. Per-tenant collections mean more ops. What do you run?
    질문답변 @David Kim · 10일 전 · 댓글 0 #rag#multi-tenant#security
  121. 5
    추천
    The agent handles maybe 60% of tickets but the 40% it hands to humans are the messy ones, so support load barely dropped. Customers who realize they're talking to a bot get annoyed even when it works. Not sure this was a win.
    자유 @Raj Patel · 3일 전 · 댓글 0 #agents#customer-support#rant
  122. 5
    추천
    We priced our product assuming $2/M input tokens. Now we're at $0.30 and dropping. Great for us, but customers keep asking why their bill didn't go down if AI got cheaper. Also, our competitors can now afford the same quality. Margin moat is shrinking.
    자유 @Marcus Obediah · 8일 전 · 댓글 0 #pricing#saas#business
  123. 5
    추천
    링크
    LangSmith is smoother but the per-seat pricing hurt us. Langfuse self-hosted costs us one small VM. Missing a few features but the trace UI is good enough and the data stays on our infra. For EU clients that matters.
    도구/프로젝트 @Tom Hall · 3일 전 · 댓글 0 #observability#langfuse#langsmith
  124. 5
    추천
    링크
    Wrote our own gateway instead of buying one. Round-robin across providers, in-memory cache for exact prompts, and a circuit breaker. It's not fancy but it cut our API bill 35% and we control everything.
    도구/프로젝트 @Raj Patel · 9일 전 · 댓글 0 #gateway#go#llm
  125. 5
    추천
    링크
    One endpoint for all providers, consistent interface, and the config is plain YAML. We swapped out three vendor SDKs in an afternoon. If you're multi-provider, just use it.
    도구/프로젝트 @Anna Nowak · 1일 전 · 댓글 0 #litellm#proxy#multivendor
  126. 5
    추천
    Like this site but for semiconductor news specifically. Scrapes 50 RSS feeds, dedupes with embeddings, summarizes with an LLM, ranks by my reading habits. It's been my morning read for a month and I've never opened the old feed reader again.
    자랑/공유 @Emily Watson · 7일 전 · 댓글 0 #news#summarization#side-project
  127. 5
    추천
    It categorizes spending, flags weird charges, and answers questions like 'how much did I spend on coffee in March'. Accuracy is 90%+ on categories. Caveat: you have to trust it with your data, so it runs fully local. Worth it for the anxiety reduction alone.
    자랑/공유 @Tom Hall · 10일 전 · 댓글 0 #finance#local-llm#privacy