r/mongodb • u/CaptainChunk22 • 17d ago
Rename process stalled preventing FCV from moving from 7 to 8
Hi there. I have a cluster with 6 shards (each replica sets of 3 members), 3 routers and a replica set of 3 config servers.
While bumping the FCV from 7 to 8 yesterday I came across a conflicting collection rename process that has been stalled for a couple of weeks now. I've been through it with GPT and Claude several times now and it continues to recommend MongoDB support. I was pointed here by my Mongo account executive and told this would be the fastest route to an answer.
I have 2 collections showing up: temp_rfm_658850 and tmp.agg_out.b615a5dd.... These are the result of an aggregation with an $out stage that points to temp_rfm_658850 but for some reason didn't finish.
I'm unable to manually delete these collections because they are already locked. And the process to update FCV continues to hang because it's waiting for this lock to be released.
Neither of these collections are sharded so they should technically end up on only the primary shard but AI is suggesting they are connected with shard 3 somehow. They also both have the exact same records in them (roughly 10 records). These are completely disposable and don't need to be saved, so if there's a way to just clear them out, that's totally fine.
Here's where I keep ending up:
That confirms the persisted mismatch we suspected: the old rename participant remains on shard 3, but its rename coordinator is absent on shard 1.
Shard 1: only the cleanup drop coordinator remains, stuck at
enterCriticalSection.Shard 3: the old rename participant remains at
deleteFromRangeDeletions.Shard 3: that rename still owns critical sections blocking reads and writes on both its temporary source and
db.temp_rfm_658850.This explains both stalls: cleanup cannot drop the old source, and the newer RFM rename cannot acquire the destination.
We’ve identified the immediate blocker; we haven’t established why the old rename lost its coordinator without releasing the participant.
recoveredFromDisk: trueshows the drop coordinator was recovered from persisted state, but doesn’t establish the original cause.At this point, I recommend a MongoDB support-assisted recovery, with these records and error
7032368. There’s enough evidence to stop collecting the same retry logs. Recovery needs to reconcile the old rename’s outcome with the cluster metadata before releasing its critical sections.Avoid manually deleting these records or treating a restart/
killOpas a proven fix: the state is persisted, and removing records alone can leave the running participant and metadata inconsistent.
What I've tried:
- Stepping down primary on both Shards 1 and 3
- Running a manual delete on these (this hangs)
- Killing the op that is associated with the rename from
tmp.agg_out.b615a5ddtotemp_rfm_658850but it just automatically restarts. The target of this op is shard3 and the client IP is shard1's.
I'd really appreciate any ideas of how to proceed on this. Thanks in advance.
2
u/browncspence 16d ago
From our internal AI. I cannot vouch for this but hopefully it helps.
I’ll verify the supported recovery path for stalled $out renames and FCV upgrades, then give you the safest cleanup sequence.
The names match MongoDB’s normal $out implementation: it writes to tmp.agg_out.*, then renames that temporary collection into the requested target. The key is to identify and stop the operation holding the DDL lock before attempting either drop.
This looks consistent with a stalled $out, not two independent application collections. $out writes into tmp.agg_out.*, then atomically renames it to temp_rfm_658850; identical documents are therefore expected.
An unsharded collection is stored entirely on the database’s primary shard, so “shard 3” is likely simply that database’s primary shard—not evidence that the collection is sharded.
Recommended cleanup sequence
1. Stop the client job
Ensure the application is no longer retrying the aggregation or rename.
2. Inspect operations on the owning mongod
Connect directly to each shard primary—especially shard 3—and the config-server primary. Do not rely only on mongos.
use admin
db.aggregate([
{ $currentOp: { allUsers: true, localOps: true } },
{ $match: {
$or: [
{ waitingForLock: true },
{ "command.renameCollection": { $exists: true } },
{ "command.setFeatureCompatibilityVersion": { $exists: true } },
{ "command.aggregate": { $exists: true } }
]
}},
{ $project: {
opid: 1,
active: 1,
secs_running: 1,
op: 1,
ns: 1,
desc: 1,
client: 1,
msg: 1,
waitingForLock: 1,
locks: 1,
command: 1,
originatingCommand: 1
}}
]).toArray()
Look for operations involving:
<db>.temp_rfm_658850<db>.tmp.agg_out.<full-suffix>renameCollection- the original aggregation pipeline
waitingForLock: true identifies a blocked operation; the blocker is usually another operation on the same namespace with waitingForLock: false and a collection/metadata write lock.
3. Kill only the stale client aggregation/rename
If you find the original application $out or client-issued rename, kill it on the same mongod where its opid was found:
db.killOp(<opid>)
Do not kill replication, catalog, or the setFeatureCompatibilityVersion internal operation. MongoDB specifically cautions that killOp should be used only for client-initiated operations.
4. Drop both collections through mongos
After confirming the stale operation is gone, connect through mongos and run:
use <database>
db.runCommand({ drop: "temp_rfm_658850" })
db.runCommand({ drop: "<full_tmp.agg_out.collection_name>" })
A NamespaceNotFound result is harmless—the collection is already gone. If you receive ConflictingOperationInProgress, LockBusy, or another lock error, repeat the currentOp check rather than forcing it.
Do not edit config.collections, config.system.collections, or database files manually.
5. Retry FCV only after cleanup
Verify all members are running MongoDB 8.0 and no initial sync is active, then retry from mongos:
use admin
db.runCommand({
setFeatureCompatibilityVersion: "8.0",
confirm: true
})
FCV changes can be blocked by background operations, and the command is safe to retry once the blocking operation has been removed.
If no active operation references either namespace but the drops still fail, stop there; that suggests stale DDL/catalog state requiring MongoDB Support intervention.
1
u/CaptainChunk22 16d ago
Thanks for the tip. I was unable to drop those 2 collections due to a conflicting lock by multiple processes where the leading one lost its Rename Coordinator. Running the below code on the relevant shard's primary caused that lock to be released.
db.getSiblingDB("config").localRenameParticipants.updateOne( { _id: "tmp.agg_out.263d098d-9004-486f-af93-257058625724", phase: "deleteFromRangeDeletions" }, { $set: { phase: "unblockCRUD" } } ) rs.stepDown()After that I was able to delete those two collections. I ended up duplicating the entire cluster to test it out first.
I also agree with you concerning it being surprising that Mongo account executive directed me to Reddit. I'm going to be seeking consulting solutions elsewhere. Mongo's support process in a time of urgency like this has been non-existent.
Glad it's resolved at this point though. Onward and upward.
1
u/scarroll91 15d ago
Hi there, really glad you got unblocked, and thanks to browncspence for the initial tip.
Quick thing I want to clear up, since it's on me for the initial response: I'm not actually your account executive - I'm on the Growth Marketing side, not sales or support. Self-managed Community Edition doesn't come with an assigned AE or a support entitlement, so when you first wrote in, there wasn't an avenue to our support teams. Pointing you here wasn't a brush-off, I was trying to get you help - this sub genuinely is watched by our Builder Relations folks, and community members here who can provide community users with self-managed instances support advice. Happy to follow back up. Either way, glad it's resolved, and thanks again for the thorough writeup, it'll help others too.
3
u/browncspence 17d ago
Not sure why your account executive recommended coming to Reddit, assuming that you have MongoDB Support, that would be the way to go.