Elasticsearch Deep Dive: Index Mapping Design and How the Inverted Index Works
Company: ByteDance
Role: Software Engineer
Category: Software Engineering Fundamentals
Difficulty: medium
Interview Round: Technical Screen
Your resume lists a backend project that uses Elasticsearch for search. The interviewer digs into the design decisions behind the project (why it was built this way, whether a different approach would have been better, and which use case each decision serves) and then into Elasticsearch itself. Answer both parts as you would for your own project.
If you are practicing without such a project, use a product-catalog search for an e-commerce site as the scenario.
### Clarifying Questions
- Which queries must the search serve: relevance-ranked keyword search, exact filters, sorting, facet counts or autocomplete? The mapping follows from them.
- Where does the data originate, and how quickly must a change in the source of truth become searchable?
- Which languages must the text analysis handle?
### Part 1 — Index and mapping design
How did you design the index and its mapping, and why? Walk through the important fields, the type and the text analysis you chose for each, and how the index is laid out and changed over time. For each decision, name the query it serves and the alternative you rejected.
```hint Start from the queries
For each field, decide whether it is searched as words, filtered or sorted as an exact value, aggregated, or only displayed; the type follows from that.
```
#### What This Part Should Cover
- Field types tied to concrete queries, including fields that need both a full-text form and an exact form
- The analyzer choice and its effect on recall and precision
- How the index is created, versioned and changed, given that an existing field's mapping cannot simply be altered
- How the index is kept in sync with the source of truth
### Part 2 — How the inverted index works
Explain the principle of Elasticsearch's inverted index: how a document's text becomes index entries, how a query is answered from them, and why this makes full-text search fast.
```hint Follow one title and one query
Take a short product title through analysis into index entries, then take a two-word query and work out which documents it finds and how.
```
#### What This Part Should Cover
- The term dictionary, and what each term's entry records about the documents that contain it
- Analysis at index time and at query time, and why the two must agree
- How new documents, updates and deletes reach the index, and why search is near-real-time rather than immediate
- What the inverted index is poor at, and which structure serves sorting and aggregations instead
### What a Strong Answer Covers
- Mapping decisions justified by specific queries rather than by defaults
- The mechanics of the inverted index, explained with a small concrete example
- Trade-offs: index size against query flexibility, and freshness against indexing cost
- A comparison with the simpler alternative of searching in the relational database
- Clear ownership of the design, including what would change if the use case changed
### Follow-up Questions
- Why is a query with a leading wildcard slow on an inverted index, and what mapping supports substring search instead?
- How do phrase queries use the information stored in the index?
- A field was mapped with the wrong type in production. How do you fix it without downtime?
- Why can the relevance score of the same document differ from one shard to another?
Overview: From a TikTok backend intern interview: explain how the Elasticsearch index and field mappings in a resume project were designed and why, then explain how an inverted index works. It tests whether a candidate understands the search technology they list, from mapping trade-offs to how queries are answered.
Read the full ByteDance Software Engineer interview experience this question came from