Steven Gonsalvez

Software Engineer


CEREBRO


machine-read, human-curated

Coding agents & tooling releases

Multi-agent & orchestration patterns

Benchmarks & eval

  • Senior SWE-Bench evaluates agents on realistic, under-specified tasks instead of junior-style over-specified ones; HN pushback: subjective LLM grading undermines rigor.
  • DeepSpec — full-stack training/eval codebase for speculative-decoding draft models.

Agentic SaaS & enterprise deployment

Data & retrieval tooling

Signals

  1. http://B.AI
  2. https://docs.b.ai/llmservice/models/claude-sonnet-5/
  3. Bloome is actually smart for coding workflows i added Claude Code, Codex, DeepSeek and a Project Lead agent into one group chat each agent…
  4. I built an MCP server so two seperate Claude Code agents can pair program without a human relaying messages
  5. deepseek-ai/DeepSpec: DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
  6. Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
  7. Kimi K2.7 Code is generally available in GitHub Copilot
  8. Autoresearch: The feedback loop behind self-improving agents
  9. v2.1.198
  10. I've been getting a TON done with Fable today and I'm not hitting rate limits. Wanted to share some tips on how I'm doing that 1. I only us…
  11. I'm working on a new skill which helps you plan enormous chunks of work, far larger than /grill-me can It identifies the frontier of decisi…
  12. RT @ClaudeDevs: Now that Fable 5 is ready to build (again), we've reset everyone's 5-hour and weekly rate limits.
  13. Proposal: a /research skill It's really simple - just spins up a background agent to look at high-trust sources, and saves them in a markdo…
  14. TencentCloud/CubeSandbox: Instant, Concurrent, Secure & Lightweight Sandbox for AI Agents.
  15. How Cursor deploys AI inside the enterprise
  16. Warp CEO Zach Lloyd on why software factories are the next phase of coding
  17. Copilot vision is generally available
  18. Browser tools for GitHub Copilot in VS Code are generally available
  19. Tabstack Browser Automation
  20. I've landed a dozen PRs and have more cooking now. Still haven't managed to hit rate limits on the $200 plan https://t.co/yo1hdDFndl
  21. Fable 5 is back. https://t.co/9RTGUCcPHy
  22. Made a tool that compiles a folder of data into something an agent can query with citations
  23. Reduced usage by using lower-effort agents when the main session is set to High or XHigh
  24. allenai/olmocr: Toolkit for linearizing PDFs for LLM datasets/training
  25. Enterprises can default to auto model selection

Sources & citations

x2099
rss1408
github803
reddit253
hackernews602
show_hn400
yc-rfs160
yc-launch30
gmail00
total57325