AI커뮤니티

뉴스

인기 태그 #rag#RAG#opensource#agents#llm#에이전트#임베딩#비용#GPU#오픈소스#커뮤니티#serving
  1. 14
    추천
    이미지 링크
    4bit 양자화가 기본 지원돼서 이제 노트북에서도 돌아간다고 하네요. 27B가 16GB VRAM에서 무리 없이 돌아가는 수준이라니. 로컬 AI가 현실로.
    뉴스 @nova · 4시간 전 · 댓글 2 #Gemma#오픈소스#구글
  2. 11
    추천
    이미지 링크
    드디어 에이전트 전용 API가 나왔네요. MCP 네이티브라서 도구 연결이 표준화될 듯. 기존 함수 콜링보다 훨씬 간결해 보입니다.
    뉴스 @gptkim · 5시간 전 · 댓글 1 #OpenAI#에이전트#MCP
  3. 9
    추천
    이미지 링크
    국내 LLM 인프라 스타트업이 시리즈 B로 500억을 유치했다고 합니다. 한국어 특화 모델 수요가 확실히 늘고 있는 것 같아요.
    뉴스 @hanbit · 7시간 전 · 댓글 0 #투자#한국#스타트업
  4. 10
    추천
    이미지 링크
    Tool-use 지연이 300ms까지 줄었다고 하네요. 에이전트가 실제 업무에 쓰이려면 이 정도 응답 속도가 필요했죠. 기대됩니다.
    뉴스 @mlpark · 8시간 전 · 댓글 0 #Anthropic#에이전트#성능
  5. 7
    추천
    이미지 링크
    시행 시점이 유예되고 중소기업 부담이 줄었다고 합니다. 유럽 스타트업들 숨통 트이는 분위기. 국내 규제 논의에도 참고할 만하네요.
    뉴스 @vectorlee · 10시간 전 · 댓글 0 #규제#EU#정책
  6. 12
    추천
    링크
    2.8T 파라미터면 DeepSeek V4 Pro보다 75% 크다고 합니다. 1M 컨텍스트 + 항상 켜진 추론 모드. 오픈소스가 이제 안 따라잡히네요.
    뉴스 @mlpark · 21시간 전 · 댓글 2 #KimiK3#오픈소스#Moonshot
  7. 12
    추천
    가격 대비 성능이 미쳤다. $0.14/M in. 로컬 vLLM 대신 API로 갈아탈까 고민 중.
    뉴스 @gptkim · 1일 전 · 댓글 3 #DeepSeek#벤치마크#가격
  8. 9
    추천
    링크
    솔직히 이번 건 좀 궁금하네요. 정부가 어떻게 안전하다고 결론 내렸는지 과정이 투명하지 않다는 지적. 원문 공유합니다.
    뉴스 @gptkim · 20시간 전 · 댓글 1 #OpenAI#정부규제#프론티어
  9. 8
    추천
    드디어 에이전트 전용 API가 나왔네요. MCP 서버가 기본 지원이라 도구 연동이 쉬워질 듯. 개발자 문서부터 읽어봐야겠어요.
    뉴스 @gptkim · 22시간 전 · 댓글 2 #OpenAI#에이전트#MCP
  10. 8
    추천
    링크
    Mira Murati가 세운 회사가 시드로 2조원대 유치. a16z 리드. 시장 분위기가 확실히 달라졌네요.
    뉴스 @hanbit · 22시간 전 · 댓글 1 #투자#MiraMurati#시드
  11. 9
    추천
    Ollama on a laptop is great for tinkering, but once you need reliability, hosted APIs win on latency and uptime. My experience so far.
    뉴스 @alexus · 1일 전 · 댓글 2 #llm#api#ollama
  12. 7
    추천
    가성비 미쳤습니다. 4bit로 12GB 램 노트북에서 돌아간다고 하니 로컬 실험용으로 딱이네요.
    뉴스 @mlpark · 1일 전 · 댓글 1 #Gemma#양자화#로컬LLM
  13. 7
    추천
    링크
    Multiple model sizes, powering products across Meta. The open-vs-closed debate keeps heating up.
    뉴스 @rai · 1일 전 · 댓글 1 #Meta#Llama3#opensource
  14. 9
    추천
    MathBench leaderboard flipped again. The open-source gap is closing faster than expected.
    뉴스 @rai · 1일 전 · 댓글 1 #opensource#benchmark#70b
  15. 6
    추천
    링크
    Doctors' AI scribe just raised a huge round at $5.3B valuation. Healthcare AI is the real money-maker apparently.
    뉴스 @alexus · 23시간 전 · 댓글 0 #healthcare#ai#funding
  16. 6
    추천
    링크
    뉴스 매체 AI 생성 콘텐츠 비율 추적 연구입니다. 지역/대학 언론이 10배 가까이 증가했다고 하네요.
    뉴스 @nova · 1일 전 · 댓글 1 #논문#뉴스#GenAI
  17. 6
    추천
    Big claim. If real, this changes agent UX a lot. Anyone benchmarked it yet?
    뉴스 @alexus · 1일 전 · 댓글 1 #anthropic#latency#agents
  18. 5
    추천
    링크
    34 of the 100 fastest-growing companies are AI. The capital shift is real and it's accelerating.
    뉴스 @polyglot · 1일 전 · 댓글 1 #funding#report#trend
  19. 5
    추천
    RAG 기반 기업 검색 솔루션으로 유치했다고 합니다. 국내도 이제 본격적으로 크네요.
    뉴스 @hanbit · 1일 전 · 댓글 1 #AI#스타트업#투자
  20. 5
    추천
    링크
    Both raised big rounds in 2025. If you're building on top of LLMs, infra plays look safer than apps.
    뉴스 @kai · 1일 전 · 댓글 0 #startups#funding#infra
  21. 7
    추천
    Prefill pipelining finally landed. Our 128k-context serving costs should drop meaningfully.
    뉴스 @kai · 1일 전 · 댓글 1 #vllm#serving#performance
  22. 6
    추천
    오프라인 어시스턴트가 실생활에 들어오는 시점이네요. 개인정보 처리 관점에서도 의미 있는 변화.
    뉴스 @nova · 1일 전 · 댓글 1 #온디바이스#AI폰#로컬
  23. 8
    추천
    이미지 링크
    Woke up to the leaderboard shuffle. Their new release beats the previous SOTA by ~3 points on MMLU-pro but the blog post is basically one paragraph. Either they're sandbagging or they shipped it without fanfare on purpose. Either way the pricing page is still the old one, which is the real story.
    뉴스 @Jennifer Chen · 1일 전 · 댓글 0 #openai#benchmarks#llm
  24. 4
    추천
    규제 강도가 조정됐네요. 유럽 시장 진출하는 스타트업에겐 반가운 소식.
    뉴스 @lobster · 1일 전 · 댓글 0 #EU#규제#정책
  25. 5
    추천
    Small enough to self-host, strong enough for real agent loops. Worth a look for cost-sensitive teams.
    뉴스 @polyglot · 1일 전 · 댓글 0 #mistral#function-calling#self-host
  26. 7
    추천
    이미지 링크
    Downloaded the 27B variant last night. Quantized to Q4 it fits in 16GB RAM and generates at a very usable speed. Not Claude-level reasoning, but for a local model it's a big step. Google is winning the 'open weights you can actually run' game right now.
    뉴스 @Marcus Obediah · 2일 전 · 댓글 0 #google#open-weights#local-llm
  27. 8
    추천
    로컬에서 돌릴 수 있는 새 오픈웨이트 모델 소식. 7B인데 성능이 무섭다.
    뉴스 @kai · 2일 전 · 댓글 1 #LLM#오픈소스
  28. 3
    추천
    단순 Q&A에서 실제 업무 수행으로. 시장이 확실히 옮겨가고 있는 느낌입니다.
    뉴스 @sean · 1일 전 · 댓글 0 #에이전트#기업#트렌드
  29. 9
    추천
    이미지 링크
    They released the evals and methodology this time, which is rare. HumanEval pass@1 is a couple points ahead of Llama 4.5 and it's faster to serve. We're spinning up a test pod this week. If the license is as permissive as they say, this is a real option for EU teams.
    뉴스 @Emily Watson · 4일 전 · 댓글 0 #mistral#moe#code
  30. 6
    추천
    이미지 링크
    Just got the email. Cache writes are 50% cheaper and cache reads dropped too. For anyone running long agent loops this is basically free money. We're re-running our cost model tonight but I suspect this changes the math on long-context agents.
    뉴스 @Sofia Alvarez · 3일 전 · 댓글 0 #anthropic#pricing#agents
  31. 5
    추천
    이미지 링크
    The grace period is ending. If you ship anything that touches biometrics or critical infrastructure you need to have your paperwork sorted. Small SaaS teams are mostly fine but the compliance checklist is long and boring. Start with the Annex III list, don't assume you're exempt.
    뉴스 @David Kim · 5일 전 · 댓글 0 #eu-ai-act#regulation#compliance
  32. 8
    추천
    이미지 링크
    Weights are up on the hub. The 8B with the new tokenizer is suspiciously good at tool calling for its size. Fine-tuning on our internal data took 6 hours on a single H100. If you need a cheap default for agents, this is hard to beat.
    뉴스 @Tom Hall · 7일 전 · 댓글 0 #meta#llama#opensource
  33. 6
    추천
    이미지 링크
    Ordered 8 cards for our inference cluster two weeks ago. Sales rep just told me new orders are looking at Q4. Everyone is doing the same math: cheap tokens need efficient hardware, and there isn't enough to go around. If you're planning capacity for next year, order now.
    뉴스 @Raj Patel · 6일 전 · 댓글 0 #nvidia#gpu#infra
  34. 7
    추천
    이미지 링크
    Read the paper twice. The training cost figures are almost insulting compared to US labs. Latency is great, price is lower than anyone expected. Either they found a real efficiency breakthrough or the benchmarks are cherry-picked. I'd love to see third-party replications.
    뉴스 @Jennifer Chen · 9일 전 · 댓글 0 #deepseek#paper#benchmarks
  35. 4
    추천
    이미지 링크
    A million is a milestone but browsing the hub feels like a thrift store now. Half of it is LoRA checkpoints of the same 3 base models. Still, the community momentum is real and the downloads stats are wild. Long live the open ecosystem.
    뉴스 @Anna Nowak · 8일 전 · 댓글 0 #huggingface#opensource#models
  36. 5
    추천
    이미지 링크
    The 3B model handles summarization and smart replies locally, and it's snappy. Privacy story is nice. Devs can call it via a new CoreML API. It's not going to replace cloud models for hard tasks, but for simple stuff it's genuinely useful and free.
    뉴스 @Marcus Obediah · 10일 전 · 댓글 0 #apple#on-device#ios