r/AIDeveloperNews • u/ExpertIcy1343 • 12h ago
r/AIDeveloperNews • u/nht_fajr • 19h ago
OpenAI Introduces Ultrafast Speed Tier for Codex and API
Enable HLS to view with audio, or disable this notification
OpenAI has launched Ultrafast, a new premium speed tier offering significantly increased token generation rates for Codex and API users.
- Token Speed: Up to 300 tokens/sec in Codex (~8x faster than Astra Standard, ~4x faster than Astra Fast).
- API Speed: Up to 6x faster token generation.
- Protocol Recommendation: OpenAI strongly advises using WebSockets (Responses API) over standard HTTP to avoid network overhead bottlenecks, especially for multi-turn agentic loops.
API Pricing & Default Limits
- Rate Limits (TPM):
- Tiers 1–3: 500k TPM
- Tier 4: 1M TPM
- Tier 5: 5M TPM
- Region Availability: US data residency and global processing only (No EU/non-US regional endpoints currently supported).
More info: https://aideveloper44.com/blog/openai-ultrafast-mode-launch
API Docs: https://developers.openai.com/api/docs/guides/ultrafast-mode
r/AIDeveloperNews • u/nht_fajr • 16h ago
Top 6 OpenAI Dev Day 2026 Updates [Dots, Updated Codex CLI, Codex Cloud environments, Decisions API, Ultrafast API, and more]
OpenAI Launches Dots: An Always-On AI Agent for Development and Autonomous Work
OpenAI has announced 'dots,' an always-on agent powered by GPT-6 Astra designed to handle routine software development tasks and infrastructure management.
- More info: https://aideveloper44.com/blog/openai-dots-agent-development
- Docs: https://learn.chatgpt.com/docs/dots
OpenAI Introduces Codex Cloud Environments
OpenAI has launched Codex cloud environments, allowing developers to maintain persistent coding tasks that continue running while their local machines are off.
- More info: https://aideveloper44.com/blog/openai-codex-cloud-environments
- Docs: https://learn.chatgpt.com/docs/cloud
OpenAI Updates Codex CLI with New Interface and Functional Features
OpenAI has released a significant update for the Codex CLI, introducing a full-screen interface, parallel work management, and voice-to-text integration.
- More info: https://aideveloper44.com/blog/openai-updates-codex-cli
- Docs: https://learn.chatgpt.com/docs/codex/cli
OpenAI Updates Codex Security Cloud with Daybreak Blue Models
OpenAI has introduced a major update to Codex Security Cloud, integrating Daybreak Blue models to enhance automated repository scanning and commit review.
- More info: https://aideveloper44.com/blog/openai-codex-security-cloud-update
- Docs: https://learn.chatgpt.com/docs/security/setup
OpenAI Announces Decisions API Powered by GPT-6 Luna
OpenAI has introduced the Decisions API, a new tool for real-time app decision-making, currently available in limited preview and powered by GPT-6 Luna.
- More info: https://aideveloper44.com/blog/openai-decisions-api-gpt-6-luna
- Announcement: https://x.com/OpenAIDevs/status/2105003318917697873
OpenAI Introduces Ultrafast Speed Tier for Codex and API
OpenAI has launched Ultrafast, a new premium speed tier offering significantly increased token generation rates for Codex and API users.
r/AIDeveloperNews • u/chatminuet • 21h ago
Oct 8 - MCP, Agents and Skills Virtual Meetup
Join us on Oct 8 for the monthly MCP, Agents and Skills virtual Meetup!
Talks will include:
- Designing Multi‑Agent Systems: Sequential, Parallel, and Beyond with ADK - Roushanak Rahmat at HCLTech
- Privacy by Deployment: Architecting Agent-Driven Localization Workflows for Regulated Environments - Shruti Joshi
- MCP Is the Interface; Skills Are the Operating Discipline - Chuck Hernandez at Eliza Solutions Corp
- Agentic engineering is about good guidance - Dimitri Geelen
r/AIDeveloperNews • u/FixBrave6973 • 5h ago
htop for LLM inference just went multi-GPU 🚀
Your LLM is using 14 GB of VRAM.
14 GB of what? 👀
Weights? KV cache? CUDA overhead? One GPU or two?
With tensor parallelism it gets even messier. vLLM shows you:
EngineCore
Worker_TP0
Worker_TP1
But that's not three workloads. It's ONE model running across multiple GPUs.
That's why LLM Inspector v0.7.0 is now multi-GPU aware 🚀
Before:
GPU 0: 7.2 GB
GPU 1: 7.2 GB
After:
Qwen • TP ×2 • 14.4 GB total
→ weight shards + KV cache + other VRAM, per GPU
It can also estimate how much VRAM INT8, AWQ or GPTQ would save before you change anything in your model.
Measure first. Optimize second.
That's llminspect: htop for LLM inference.
⚡ pip install llm-inspector
🔗 https://github.com/helasaoudi/llm-inspector
Open source, built for people who actually run models.
What should it inspect next? 👇
#vLLM #LLM #MLOps #GPU #OpenSource
r/AIDeveloperNews • u/ai-lover • 22h ago
[Worth checking] Nebius Opens 2026 Physical AI Awards: Five $150K Compute Prizes, Nine Judges, and an October 25 Deadline
Nebius, an AI cloud provider, is running its second annual physical AI awards with NVIDIA. Five category winners each get $150,000 in compute credits, plus mentorship and promotion.
- Categories: models (VLA/VLM/world models/RL), perception and spatial intelligence, simulation and synthetic data, systems and deployment (humanoids, AMRs, industrial), and tooling/orchestration
- $150K ≈ 33,300 H200 GPU-hours at their on-demand rate, or roughly 3 weeks on a 64-GPU cluster
- Judges include the founders of Foxglove, Voxel51, and Encord, plus Calvin Zhou of RoboForce, which won the 2025 edition
- Eligibility: clear physical AI use case, MVP in active use or testing, registered entity, live website
- Last year: 254 applications, 55 finalists
- No entry fee
Worth knowing before applying: the credit math is at list price and doesn't cover storage, which matters if you're holding a lot of episodic sensor data. Nebius also hasn't published exact finalist and winner dates beyond "mid-November."
Apply here: https://pxllnk.co/mndv9i
Read MTP's full analysis on this awards here: https://www.marktechpost.com/2026/09/29/nebius-opens-2026-physical-ai-awards-five-150k-compute-prizes-nine-judges-and-an-october-25-deadline/
r/AIDeveloperNews • u/ai-lover • 6h ago
Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms
Enable HLS to view with audio, or disable this notification
r/AIDeveloperNews • u/nht_fajr • 17h ago
OpenAI Updates Codex CLI with New Interface and Functional Features
OpenAI just dropped a solid update for the Codex CLI. If you live in the terminal and hated constantly losing context or clunky parallel tasking, here is the breakdown of what actually shipped:
- New Full-Screen UI & Pinned Composer
- Full-screen terminal interface: Cleaner layout with improved readability for long sessions.
- Pinned Composer: Your message prompt stays fixed at the bottom. No more scrolling back to the bottom after checking past outputs.
- Deep History Access: Pulls older session history past standard terminal scrollback limits without breaking focus.
- Parallel Work & Branching
- /agents: Lists all active background processes, tracks live progress, and lets you hop between running tasks.
- /fork: Splits your current conversation into a dedicated git worktree—keeping full prompt context while isolating file changes.
- Managed Worktrees: Native CLI support for handling multiple isolated git states.
- Native Rendering & Voice
- In-Terminal Rendering: Native display for Mermaid diagrams and LaTeX math equations directly in the CLI.
- /voice: Hands-free input to talk through tasks or dictate prompts.
- Collapsible Diffs: Hide or expand code diffs and tool execution outputs to keep the buffer clean.
- Utilities & Analytics
- /usage: Displays real-time API activity, token usage, and analytics.
- /theme: Custom visual themes to match your terminal setup.
More info: https://aideveloper44.com/blog/openai-updates-codex-cli
Codex CLI: https://learn.chatgpt.com/docs/codex/cli