Infrastructure and model selection
I see teams attempting to deploy Cohere Command R+ using deprecated OCI models or legacy generic shapes. The OCI Generative AI service deprecated the Cohere Command R (cohere.command-r-16k) and Command R+ (cohere.command-r-plus) models. You must use the hardware unit shapes for new dedicated AI clusters. If you use a dedicated AI cluster with a legacy generic Cohere shape during the retirement period, you might see both the legacy generic shapes and the new hardware unit shapes in the API. For the Command A model in UAE East (Dubai), you must multiply the unit price by 4. For the standard Command A, you need 1 LARGE_COHERE_V3 unit. The Command R 08-2024 model provides a 128,000 token context length and supports fine-tuning with LoRA and T-Few training methods. This model is available in US Midwest (Chicago), Germany Central (Frankfurt), UK South (London), and Brazil East (Sao Paulo). Command R 08-2024 rivals the previous R+ version in math, coding, and reasoning. Command A has a 256,000 token context window.
| Model Name | OCI Model Name | Unit Size | Required Units |
|---|---|---|---|
| Cohere Command A | cohere.command-a-03-2025 | LARGE_COHERE_V3 | 1 |
| Command A (UAE East) | cohere.command-a-03-2025 | SMALL_COHERE_4 | 1 |
Retrieval pitfalls
Retrieval precision suffers when teams ignore document structure or content age. A document re-publication creates a freshness gap if the vector store holds old chunks. The retriever scores by semantic similarity instead of ingestion timestamp, so the old chunk wins. You must add a recency multiplier to the relevance score before the final top-k cut. You should also use structure-aware parsers like docling to segment PDFs into typed elements like headings, paragraphs, and tables. If a PDF parser flattens a table into a row of pipe characters, the embedder sees a bag of cell values rather than column-row relationships. Most teams skip the parent-document retrieval method that matches small clauses to large sections. You should implement hierarchical chunking. You can also encounter failure when the LLM fails to extract the answer from the retrieved context due to noise or ambiguity. If the model provides the answer in the wrong format, use structured outputs like JSON. You know that a large context window does not replace a retrieval pipeline. I would avoid the error of using a single embedder without implementing hierarchical chunking.
The deployment wall
I observe that 70% to 80% of enterprise RAG projects fail before they reach production. Many teams copy a tutorial that wires a single embedder, a single vector store, and a single LLM call. This architecture fails when requirements change. A production-ready system requires a structured ingestion layer, a chunking layer with metadata, a hybrid retriever, a reranker, a freshness filter, an LLM with cited generation, and an evaluation loop to prevent the typical failures seen in pilot projects. You must implement these seven parts or you will face failures in the third week of the rollout. The 2026 production RAG architecture write-ups suggest that a working system has seven moving parts. These parts include a structured ingestion layer, a chunking layer with metadata, a hybrid retriever, a reranker, a freshness filter, an LLM with cited generation, and an evaluation loop running async on production traffic. None of those are optional. All of them are what teams skip in the demo and pay for in week three of the rollout. Roughly 70% of teams running RAG in production have no systematic evaluation for retrieval quality according to Ragaboutit. They might use an LLM judge for end-to-end answer quality. However, they lack metrics for context recall and context precision. Do you have a 200-question gold set to run on every embedding change? Ragas provides these metrics for free. Barnett et al. (2024) identified seven failure points, including missing content and missing top-ranked documents. If the correct information exists in your dataset but the ranker places it too low, the model misses the answer.
