# How to identify and delete invalid child IDs

**URL:** <https://discuss.elastic.co/t/how-to-identify-and-delete-invalid-child-ids/321512>\
**Category:** Elasticsearch\
**Created:** [December 19, 2022, 2:57am UTC](https://discuss.elastic.co/t/how-to-identify-and-delete-invalid-child-ids/321512 "2022-12-19T02:57:48Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![kgeographer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kgeographer/32/26217_2.png) [@kgeographer](https://discuss.elastic.co/u/kgeographer)\
**Post date:** [December 19, 2022, 2:57am UTC](https://discuss.elastic.co/t/how-to-identify-and-delete-invalid-child-ids/321512/1 "2022-12-19T02:57:48Z")

</div>

I have an index making extensive use of Parent/Child relations, and over time some Child docs have been deleted without removing reference to them in their Parent. This doesn't impact searches but I do use a count of Children in ordering, so I need to prune these 'zombie' references to non-existing Children from Parents.

I can imagine a brute force approach retrieving each Parent having Children, then getting all Child docs and doing a pythonic set comparison of ids, but is there a more efficient approach?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 19, 2022, 5:52am UTC](https://discuss.elastic.co/t/how-to-identify-and-delete-invalid-child-ids/321512/2 "2022-12-19T05:52:28Z")

</div>

> [@kgeographer](#):
>
> without removing reference to them in their Parent

Typically parent-child relationships do not have references to children in the parent as far as I know, so is this something you have added and maitain?

---

<div class="post-metadata">

**Author:** ![kgeographer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kgeographer/32/26217_2.png) [@kgeographer](https://discuss.elastic.co/u/kgeographer)\
**Post date:** [December 19, 2022, 12:26pm UTC](https://discuss.elastic.co/t/how-to-identify-and-delete-invalid-child-ids/321512/3 "2022-12-19T12:26:30Z")

</div>

Thanks for the reminder! Built this 5 years ago.

Yes, I have added a children array field, and each time a child is assigned an exiting parent, its \_id gets added to the parent's `children[]`. There's also a `searchy[]` array, and certain tag terms of the child get added to that.

This seemed at the time (and now) convoluted, but it was the only way I could think of to meet my requirement: searching for a tag and returning one or more parent/child 'clusters'. That is all of the parent+children **sets** that have a given tag in any of their members, whether parent or child. So the search for tags is limited to parents, and returns the parent and its child \_ids. It approximates a graph in a way.

Maintenance is proving to be a struggle, because over time documents can get updated or removed by a web app, meaning surgery: removing a parent that has children requires transferring the parent role to one of them. Removing a child requires removing its \_id from the parents `children[]` field...not to mention maintenance of the `searchy[]` field.

Seems I've dug myself a hole - suggestions are welcome!

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 19, 2022, 12:29pm UTC](https://discuss.elastic.co/t/how-to-identify-and-delete-invalid-child-ids/321512/4 "2022-12-19T12:29:29Z")

</div>

> [@kgeographer](#):
>
> I can imagine a brute force approach retrieving each Parent having Children, then getting all Child docs and doing a pythonic set comparison of ids, but is there a more efficient approach?

I can not see any easy or efficient other way to get around this.

---

<div class="post-metadata">

**Author:** ![kgeographer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kgeographer/32/26217_2.png) [@kgeographer](https://discuss.elastic.co/u/kgeographer)\
**Post date:** [December 19, 2022, 12:36pm UTC](https://discuss.elastic.co/t/how-to-identify-and-delete-invalid-child-ids/321512/5 "2022-12-19T12:36:44Z")

</div>

Thanks. I plan to investigate redesigning this architecture, as the nature of the data is in fact graph-like. That research is in front of me: [Graph: Explore Connections in Elasticsearch Data | Elastic](https://www.elastic.co/what-is/elasticsearch-graph)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 16, 2023, 12:37pm UTC](https://discuss.elastic.co/t/how-to-identify-and-delete-invalid-child-ids/321512/6 "2023-01-16T12:37:28Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
