AICommunity

All

Popular tags #rag#RAG#opensource#agents#llm#에이전트#임베딩#비용#GPU#오픈소스#커뮤니티#serving
  1. 7
    votes
    Link
    Updated list of open LLMs you can use commercially. Handy when picking a model for a client project.
    Tools/Projects @sean · 21시간 전 · comments 1 #opensource#llm#list
  2. 9
    votes
    Ollama on a laptop is great for tinkering, but once you need reliability, hosted APIs win on latency and uptime. My experience so far.
    News @alexus · 1일 전 · comments 2 #llm#api#ollama
  3. 8
    votes
    Every framework becomes a debugging nightmare past 500 lines. Plain code + function calling has been better for us.
    Free @polyglot · 1일 전 · comments 1 #agents#frameworks#opinion
  4. 7
    votes
    All data stays on device. Weekly digest with topics, mood, and action items. 400 users in 2 weeks!
    Showcase @kai · 1일 전 · comments 1 #local-first#journal#product
  5. 7
    votes
    Link
    Multiple model sizes, powering products across Meta. The open-vs-closed debate keeps heating up.
    News @rai · 1일 전 · comments 1 #Meta#Llama3#opensource
  6. 9
    votes
    MathBench leaderboard flipped again. The open-source gap is closing faster than expected.
    News @rai · 1일 전 · comments 1 #opensource#benchmark#70b
  7. 6
    votes
    Link
    Doctors' AI scribe just raised a huge round at $5.3B valuation. Healthcare AI is the real money-maker apparently.
    News @alexus · 23시간 전 · comments 0 #healthcare#ai#funding
  8. 9
    votes
    Ticket deflection went from 0% to 38% in a month. Here's the architecture and the mistakes we made.
    Showcase @alexus · 1일 전 · comments 1 #support#rag#case-study
  9. 6
    votes
    We went LibreChat for multi-user + RBAC. Open-webui has nicer UX though. Your pick?
    Tools/Projects @kai · 1일 전 · comments 1 #self-host#chat-ui#open-source
  10. 7
    votes
    Mine: Claude for code, DeepSeek for bulk tasks, local Ollama for privacy stuff. What's yours?
    Free @alexus · 1일 전 · comments 1 #stack#daily#workflow
  11. 5
    votes
    We use RAGAS but the scores don't correlate with human judgment. What do you actually use in production?
    Q&A @polyglot · 1일 전 · comments 1 #rag#evaluation#ragas
  12. 6
    votes
    Scrapes 20 sources, dedupes, summarizes, posts at 8am. 1.2k subscribers now. AMA about the pipeline.
    Showcase @rai · 1일 전 · comments 1 #telegram#bot#newsletter
  13. 6
    votes
    Big claim. If real, this changes agent UX a lot. Anyone benchmarked it yet?
    News @alexus · 1일 전 · comments 1 #anthropic#latency#agents
  14. 5
    votes
    Link
    Keep this tab open. Everything from o1 to GPT-5-Codex-Mini lands here first.
    Tools/Projects @alexus · 1일 전 · comments 1 #openai#changelog#models
  15. 5
    votes
    Link
    34 of the 100 fastest-growing companies are AI. The capital shift is real and it's accelerating.
    News @polyglot · 1일 전 · comments 1 #funding#report#trend
  16. 6
    votes
    Chroma vs Qdrant vs LanceDB for a small self-hosted RAG service. Disk and RAM matter — running on a free-tier VM.
    Q&A @kai · 1일 전 · comments 2 #vectordb#rag
  17. 7
    votes
    Ran a quick benchmark. pgvector with HNSW was 80% the speed at 10% the cost. Details in thread.
    Tools/Projects @polyglot · 1일 전 · comments 0 #pgvector#pinecone#benchmark
  18. 5
    votes
    Link
    Both raised big rounds in 2025. If you're building on top of LLMs, infra plays look safer than apps.
    News @kai · 1일 전 · comments 0 #startups#funding#infra
  19. 7
    votes
    Feed it a CSV export, it finds anomalies and writes a Korean explanation via an LLM. Useful for MSP folks.
    Tools/Projects @sean · 1일 전 · comments 1 #cloud#cost#cli
  20. 7
    votes
    Prefill pipelining finally landed. Our 128k-context serving costs should drop meaningfully.
    News @kai · 1일 전 · comments 1 #vllm#serving#performance
  21. 5
    votes
    Link
    One endpoint for all providers, consistent interface, and the config is plain YAML. We swapped out three vendor SDKs in an afternoon. If you're multi-provider, just use it.
    Tools/Projects @Anna Nowak · 1일 전 · comments 0 #litellm#proxy#multivendor
  22. 6
    votes
    Long agent sessions blow up tokens. Do you use summarization, sliding windows, or something else?
    Q&A @kai · 1일 전 · comments 1 #context#agents#tokens
  23. 5
    votes
    I packaged my go-to RAG setup into a template. Docker compose up and you have a working API in 2 minutes.
    Tools/Projects @sean · 1일 전 · comments 1 #docker#ollama#fastapi#template
  24. 8
    votes
    Replaced a custom fine-tune with a generic 8B + good retrieval. Cheaper to maintain, better scores. Data point for you.
    Showcase @vectorlee · 1일 전 · comments 1 #rag#finetune#case-study
  25. 4
    votes
    Link
    Worth a read — some genuinely interesting verticals (last-mile routing, predictive maintenance).
    Showcase @rai · 1일 전 · comments 1 #startups#ml#vertical
  26. 8
    votes
    Image Link
    Woke up to the leaderboard shuffle. Their new release beats the previous SOTA by ~3 points on MMLU-pro but the blog post is basically one paragraph. Either they're sandbagging or they shipped it without fanfare on purpose. Either way the pricing page is still the old one, which is the real story.
    News @Jennifer Chen · 1일 전 · comments 0 #openai#benchmarks#llm
  27. 6
    votes
    It joins the call with a normal name, takes notes, and posts them to our wiki. People forget it's there until the notes appear. Uses a tiny local model for transcription. The 'is this legal' question came up once. HR said fine, so we're shipping it.
    Showcase @Marcus Obediah · 1일 전 · comments 0 #meetings#notes#side-project
  28. 6
    votes
    Model cites chunks that don't actually support the answer. Any post-processing tricks that work?
    Q&A @rai · 1일 전 · comments 0 #rag#hallucination#citations
  29. 5
    votes
    vLLM vs TGI vs llama.cpp server. For a small team serving one 8B model, what's the least ops burden?
    Q&A @alexus · 1일 전 · comments 1 #self-host#serving#llm
  30. 5
    votes
    Visual node editor for agent pipelines with live tracing. MIT licensed, would love feedback.
    Showcase @polyglot · 1일 전 · comments 0 #opensource#agents#playground
  31. 6
    votes
    CSV of prompts → JSONL of responses + scores. 100 lines of Go. Maybe useful for your eval pipeline.
    Tools/Projects @alexus · 1일 전 · comments 0 #cli#evals#opensource
  32. 5
    votes
    Gradio got multimodal chat components. Streamlit is catching up. For internal demos, which wins?
    Tools/Projects @rai · 1일 전 · comments 1 #gradio#streamlit#demo
  33. 6
    votes
    With better rerankers and hybrid search, raw embedding quality matters less every quarter. Agree?
    Free @sean · 1일 전 · comments 0 #embeddings#trend#hot-take
  34. 5
    votes
    I used to read every arxiv paper. Now I wait for good summaries. Am I falling behind?
    Free @rai · 1일 전 · comments 0 #papers#learning#arxiv
  35. 5
    votes
    Small enough to self-host, strong enough for real agent loops. Worth a look for cost-sensitive teams.
    News @polyglot · 1일 전 · comments 0 #mistral#function-calling#self-host
  36. 6
    votes
    Scrapes arXiv + HN, summarizes with an LLM, posts as threads. 200 stars in a week. AMA.
    Showcase @rai · 1일 전 · comments 2 #newsletter#automation
  37. 5
    votes
    I keep coming back to plain code + function calling. Frameworks add too much magic for debugging.
    Free @polyglot · 1일 전 · comments 1 #agents#function-calling
  38. 4
    votes
    Lots of 'prompt engineer' listings that are really content farms. Anyone found legit remote AI work lately?
    Free @kai · 1일 전 · comments 1 #remote#work#jobs
  39. 7
    votes
    Image Link
    Downloaded the 27B variant last night. Quantized to Q4 it fits in 16GB RAM and generates at a very usable speed. Not Claude-level reasoning, but for a local model it's a big step. Google is winning the 'open weights you can actually run' game right now.
    News @Marcus Obediah · 2일 전 · comments 0 #google#open-weights#local-llm
  40. 4
    votes
    This week alone: three new models, two frameworks, one protocol update. I've started ignoring anything that isn't usable in my stack today. My feed is a graveyard of 'this changes everything' posts I never read. What's your filter?
    Free @Raj Patel · 1일 전 · comments 0 #news#overwhelm#community
  41. 6
    votes
    Ship a model, top the leaderboard, get the press release, repeat. Meanwhile real usage is full of edge cases none of the benchmarks touch. I'd rather have a model that's honest about what it can't do.
    Free @Emily Watson · 2일 전 · comments 0 #benchmarks#opinion#hot-take
  42. 8
    votes
    Our org has 10k pages of policy PDFs nobody reads. Built a RAG bot with citations that links back to the exact page. Support queries dropped noticeably and compliance is happy because everything is traceable. Side project that became the team's favorite tool.
    Showcase @Jennifer Chen · 2일 전 · comments 0 #rag#chatbot#showcase
  43. 4
    votes
    System prompt first, stable ordering, minimal drift. Any other tricks to keep cache hits high?
    Q&A @polyglot · 2일 전 · comments 1 #prompt#cache#llm
  44. 3
    votes
    We need to ground answers in internal docs that update weekly. RAG feels obvious but retrieval quality is killing us. Fine-tuning is a one-time cost but the docs change. Anyone run a hybrid? What's your split?
    Q&A @Sofia Alvarez · 1일 전 · comments 0 #rag#finetuning#grounding
  45. 6
    votes
    Link
    We were merging PRs blind. Now every RAG change runs 40 questions with expected answers and fails the build if recall drops. It's crude but it caught two regressions last week. Sometimes the simple stuff wins.
    Tools/Projects @Anna Nowak · 2일 전 · comments 0 #rag#evals#testing
  46. 9
    votes
    Image Link
    They released the evals and methodology this time, which is rare. HumanEval pass@1 is a couple points ahead of Llama 4.5 and it's faster to serve. We're spinning up a test pod this week. If the license is as permissive as they say, this is a real option for EU teams.
    News @Emily Watson · 4일 전 · comments 0 #mistral#moe#code
  47. 6
    votes
    Image Link
    Just got the email. Cache writes are 50% cheaper and cache reads dropped too. For anyone running long agent loops this is basically free money. We're re-running our cost model tonight but I suspect this changes the math on long-context agents.
    News @Sofia Alvarez · 3일 전 · comments 0 #anthropic#pricing#agents
  48. 4
    votes
    Long context is expensive and my tests show the model forgets the middle anyway. Chunking into a RAG loop works but feels like giving up. Do you just pay for big context or actually engineer around it?
    Q&A @David Kim · 2일 전 · comments 0 #context#rag#llm
  49. 5
    votes
    The agent handles maybe 60% of tickets but the 40% it hands to humans are the messy ones, so support load barely dropped. Customers who realize they're talking to a bot get annoyed even when it works. Not sure this was a win.
    Free @Raj Patel · 3일 전 · comments 0 #agents#customer-support#rant
  50. 7
    votes
    Link
    50 lines of TypeScript, points at our search API, and now everyone queries docs from their editor. It's dumb but it changed how we work. The protocol is clunky but the integration win is real.
    Tools/Projects @Marcus Obediah · 4일 전 · comments 0 #mcp#internal-tools#docs