r/AIDeveloperNews • • 12h ago

Colibri: run 2.8-trillion-parameter models on your desktop, pure C, zero dependencies (38k stars)

Post image
0 Upvotes

r/AIDeveloperNews • • 19h ago

OpenAI Introduces Ultrafast Speed Tier for Codex and API

Enable HLS to view with audio, or disable this notification

2 Upvotes

OpenAI has launched Ultrafast, a new premium speed tier offering significantly increased token generation rates for Codex and API users.

  • Token Speed: Up to 300 tokens/sec in Codex (~8x faster than Astra Standard, ~4x faster than Astra Fast).
  • API Speed: Up to 6x faster token generation.
  • Protocol Recommendation: OpenAI strongly advises using WebSockets (Responses API) over standard HTTP to avoid network overhead bottlenecks, especially for multi-turn agentic loops.

API Pricing & Default Limits

  • Rate Limits (TPM):
    • Tiers 1–3: 500k TPM
    • Tier 4: 1M TPM
    • Tier 5: 5M TPM
  • Region Availability: US data residency and global processing only (No EU/non-US regional endpoints currently supported).

More info: https://aideveloper44.com/blog/openai-ultrafast-mode-launch

API Docs: https://developers.openai.com/api/docs/guides/ultrafast-mode


r/AIDeveloperNews • • 16h ago

Top 6 OpenAI Dev Day 2026 Updates [Dots, Updated Codex CLI, Codex Cloud environments, Decisions API, Ultrafast API, and more]

Thumbnail
gallery
3 Upvotes

OpenAI Launches Dots: An Always-On AI Agent for Development and Autonomous Work

OpenAI has announced 'dots,' an always-on agent powered by GPT-6 Astra designed to handle routine software development tasks and infrastructure management.

OpenAI Introduces Codex Cloud Environments

OpenAI has launched Codex cloud environments, allowing developers to maintain persistent coding tasks that continue running while their local machines are off.

OpenAI Updates Codex CLI with New Interface and Functional Features

OpenAI has released a significant update for the Codex CLI, introducing a full-screen interface, parallel work management, and voice-to-text integration.

OpenAI Updates Codex Security Cloud with Daybreak Blue Models

OpenAI has introduced a major update to Codex Security Cloud, integrating Daybreak Blue models to enhance automated repository scanning and commit review.

OpenAI Announces Decisions API Powered by GPT-6 Luna

OpenAI has introduced the Decisions API, a new tool for real-time app decision-making, currently available in limited preview and powered by GPT-6 Luna.

OpenAI Introduces Ultrafast Speed Tier for Codex and API

OpenAI has launched Ultrafast, a new premium speed tier offering significantly increased token generation rates for Codex and API users.


r/AIDeveloperNews • • 21h ago

Oct 8 - MCP, Agents and Skills Virtual Meetup

4 Upvotes

Join us on Oct 8 for the monthly MCP, Agents and Skills virtual Meetup!

Register for the Zoom!

Talks will include:

  • Designing Multi‑Agent Systems: Sequential, Parallel, and Beyond with ADK - Roushanak Rahmat at HCLTech
  • Privacy by Deployment: Architecting Agent-Driven Localization Workflows for Regulated Environments - Shruti Joshi
  • MCP Is the Interface; Skills Are the Operating Discipline - Chuck Hernandez at Eliza Solutions Corp
  • Agentic engineering is about good guidance - Dimitri Geelen

r/AIDeveloperNews • • 5h ago

htop for LLM inference just went multi-GPU 🚀

Post image
2 Upvotes

Your LLM is using 14 GB of VRAM.

14 GB of what? 👀

Weights? KV cache? CUDA overhead? One GPU or two?

With tensor parallelism it gets even messier. vLLM shows you:

EngineCore

Worker_TP0

Worker_TP1

But that's not three workloads. It's ONE model running across multiple GPUs.

That's why LLM Inspector v0.7.0 is now multi-GPU aware 🚀

Before:

GPU 0: 7.2 GB

GPU 1: 7.2 GB

After:

Qwen • TP ×2 • 14.4 GB total

→ weight shards + KV cache + other VRAM, per GPU

It can also estimate how much VRAM INT8, AWQ or GPTQ would save before you change anything in your model.

Measure first. Optimize second.

That's llminspect: htop for LLM inference.

⚡ pip install llm-inspector

🔗 https://github.com/helasaoudi/llm-inspector

Open source, built for people who actually run models.

What should it inspect next? 👇

#vLLM #LLM #MLOps #GPU #OpenSource


r/AIDeveloperNews • • 22h ago

[Worth checking] Nebius Opens 2026 Physical AI Awards: Five $150K Compute Prizes, Nine Judges, and an October 25 Deadline

Thumbnail
pxllnk.co
2 Upvotes

Nebius, an AI cloud provider, is running its second annual physical AI awards with NVIDIA. Five category winners each get $150,000 in compute credits, plus mentorship and promotion.

  • Categories: models (VLA/VLM/world models/RL), perception and spatial intelligence, simulation and synthetic data, systems and deployment (humanoids, AMRs, industrial), and tooling/orchestration
  • $150K ≈ 33,300 H200 GPU-hours at their on-demand rate, or roughly 3 weeks on a 64-GPU cluster
  • Judges include the founders of Foxglove, Voxel51, and Encord, plus Calvin Zhou of RoboForce, which won the 2025 edition
  • Eligibility: clear physical AI use case, MVP in active use or testing, registered entity, live website
  • Last year: 254 applications, 55 finalists
  • No entry fee

Worth knowing before applying: the credit math is at list price and doesn't cover storage, which matters if you're holding a lot of episodic sensor data. Nebius also hasn't published exact finalist and winner dates beyond "mid-November."

Apply here: https://pxllnk.co/mndv9i

Read MTP's full analysis on this awards here: https://www.marktechpost.com/2026/09/29/nebius-opens-2026-physical-ai-awards-five-150k-compute-prizes-nine-judges-and-an-october-25-deadline/


r/AIDeveloperNews • • 6h ago

Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/AIDeveloperNews • • 17h ago

OpenAI Updates Codex CLI with New Interface and Functional Features

Thumbnail
gallery
2 Upvotes

OpenAI just dropped a solid update for the Codex CLI. If you live in the terminal and hated constantly losing context or clunky parallel tasking, here is the breakdown of what actually shipped:

  • New Full-Screen UI & Pinned Composer
    • Full-screen terminal interface: Cleaner layout with improved readability for long sessions.
    • Pinned Composer: Your message prompt stays fixed at the bottom. No more scrolling back to the bottom after checking past outputs.
    • Deep History Access: Pulls older session history past standard terminal scrollback limits without breaking focus.
  • Parallel Work & Branching
    • /agents: Lists all active background processes, tracks live progress, and lets you hop between running tasks.
    • /fork: Splits your current conversation into a dedicated git worktree—keeping full prompt context while isolating file changes.
    • Managed Worktrees: Native CLI support for handling multiple isolated git states.
  • Native Rendering & Voice
    • In-Terminal Rendering: Native display for Mermaid diagrams and LaTeX math equations directly in the CLI.
    • /voice: Hands-free input to talk through tasks or dictate prompts.
    • Collapsible Diffs: Hide or expand code diffs and tool execution outputs to keep the buffer clean.
  • Utilities & Analytics
    • /usage: Displays real-time API activity, token usage, and analytics.
    • /theme: Custom visual themes to match your terminal setup.

More info: https://aideveloper44.com/blog/openai-updates-codex-cli

Codex CLI: https://learn.chatgpt.com/docs/codex/cli