Prompting an LLM to Migrate a Java Codebase to Python and Clean a Large Database

Read the full interview experience this question came from →

Quick Overview

Explain how you would prompt a large language model to migrate a Java codebase to Python and to help clean a database with tens of thousands of rows. It tests prompt design, verification of generated code and data fixes, data quality measurement, security of AI-assisted work and the handling of sensitive data.

Prompting an LLM to Migrate a Java Codebase to Python and Clean a Large Database

Company: Disney

Role: Data Engineer

Category: Software Engineering Fundamentals

Difficulty: easy

Interview Round: Onsite

Your team wants to use a large language model (LLM) to speed up two pieces of data engineering work: 1. migrating an entire codebase from Java to Python, and 2. cleaning a database with tens of thousands of rows of inconsistent data. Explain how you would prompt the LLM for each task, and how you would ensure data quality and security throughout, including how you would deal with sensitive data. ### Constraints and Clarifications - The model, the tooling (chat interface, IDE assistant, API) and the contents of the codebase and the database are not specified. State your assumptions. - The focus is on how you instruct the model and on the safeguards around it, not on any one product. ### Clarifying Questions - Is the LLM a public service, an enterprise deployment with contractual limits on data retention and training, or a self-hosted model? - How large is the Java codebase, which frameworks and libraries does it use, and how good is its test coverage? - What does "clean" mean here: duplicates, inconsistent formats, invalid or missing values, records that need categorizing? - Does the database hold personal or otherwise regulated data? ### Part 1 — Prompting a Java-to-Python migration How would you prompt the LLM to migrate the whole codebase from Java to Python? Explain how you split the work, what context each prompt carries, and how you establish that the result is correct. ```hint One prompt will not do A whole codebase is more than a model can reliably translate in one pass. Think about the order in which to migrate modules, and what each prompt must know about the code around it. ``` ```hint Define done first Decide what evidence would convince you that a translated module behaves exactly like the Java original. ``` #### What This Part Should Cover - Splitting the work by module and dependency order, with shared conventions in every prompt - Explicit rules for Java constructs and libraries whose behavior differs in Python - Verification against the original: tests, recorded outputs, side-by-side runs, review - An iteration loop on failures, with a human accountable for every merge ### Part 2 — Prompting a database cleanup How would you prompt the LLM to help clean a database with tens of thousands of rows? ```hint Rules or rows Consider whether the model should edit each row itself, or produce something you can review once and then apply to every row. ``` ```hint Rows that fit no rule Some values will be ambiguous. Decide where they go instead of letting the model guess. ``` #### What This Part Should Cover - Profiling the data first, and what the model is shown: schema, statistics, samples - Model-proposed rules turned into reviewed code, versus per-row model calls for genuinely fuzzy cases - Structured, validated output and a review path for uncertain records - Non-destructive execution with an audit trail and rollback ### Part 3 — Data quality, security and sensitive data How do you make sure the migrated code and the cleaned data are of high quality and that the work is secure? How do you deal with sensitive data along the way? ```hint What leaves your boundary Trace which code and which data each prompt sends, where it goes, and whether it is retained there. ``` ```hint Prove it improved Choose measurements, taken before and after the cleanup, that would show quality improved and nothing was lost. ``` #### What This Part Should Cover - Measurable quality checks before and after, and reconciliation of what changed - Where the model runs and what data it may see: minimization, masking, synthetic samples - Security review of generated code, access control, and auditing of prompts and outputs - Compliance obligations for regulated data ### What a Strong Answer Covers - Treating the LLM as a fast but fallible assistant whose output is verified, not trusted - A concrete prompt structure: role, context, constraints, examples and output format - Reviewable, deterministic artifacts (code, SQL, rules) instead of unexplained bulk edits - Quality that is measured, with changes that can be reversed - Sensitive data kept out of prompts unless the deployment is approved for it ### Follow-up Questions - The translated Python passes every test but runs much slower than the Java version. How do you find out why, and would you ask the LLM to fix it? - What would it cost in time and tokens to send every row through the LLM, compared with having it write cleanup rules? - The model "corrects" a value that was actually right. How does your process catch it? - A teammate pasted production customer records into a public chatbot. What do you do now, and what prevents it next time?

Overview: Explain how you would prompt a large language model to migrate a Java codebase to Python and to help clean a database with tens of thousands of rows. It tests prompt design, verification of generated code and data fixes, data quality measurement, security of AI-assisted work and the handling of sensitive data.

Read the full Disney Data Engineer interview experience this question came from

|Home/Software Engineering Fundamentals/Disney
Disney logo
Disney
Aug 31, 2026
easyData EngineerOnsiteSoftware Engineering Fundamentals
0
0

Your team wants to use a large language model (LLM) to speed up two pieces of data engineering work:

  1. migrating an entire codebase from Java to Python, and
  2. cleaning a database with tens of thousands of rows of inconsistent data.

Explain how you would prompt the LLM for each task, and how you would ensure data quality and security throughout, including how you would deal with sensitive data.

Constraints and Clarifications

  • The model, the tooling (chat interface, IDE assistant, API) and the contents of the codebase and the database are not specified. State your assumptions.
  • The focus is on how you instruct the model and on the safeguards around it, not on any one product.

Clarifying Questions Guidance

  • Is the LLM a public service, an enterprise deployment with contractual limits on data retention and training, or a self-hosted model?
  • How large is the Java codebase, which frameworks and libraries does it use, and how good is its test coverage?
  • What does "clean" mean here: duplicates, inconsistent formats, invalid or missing values, records that need categorizing?
  • Does the database hold personal or otherwise regulated data?

Part 1 — Prompting a Java-to-Python migration

How would you prompt the LLM to migrate the whole codebase from Java to Python? Explain how you split the work, what context each prompt carries, and how you establish that the result is correct.

What This Part Should Cover Guidance

  • Splitting the work by module and dependency order, with shared conventions in every prompt
  • Explicit rules for Java constructs and libraries whose behavior differs in Python
  • Verification against the original: tests, recorded outputs, side-by-side runs, review
  • An iteration loop on failures, with a human accountable for every merge

Part 2 — Prompting a database cleanup

How would you prompt the LLM to help clean a database with tens of thousands of rows?

What This Part Should Cover Guidance

  • Profiling the data first, and what the model is shown: schema, statistics, samples
  • Model-proposed rules turned into reviewed code, versus per-row model calls for genuinely fuzzy cases
  • Structured, validated output and a review path for uncertain records
  • Non-destructive execution with an audit trail and rollback

Part 3 — Data quality, security and sensitive data

How do you make sure the migrated code and the cleaned data are of high quality and that the work is secure? How do you deal with sensitive data along the way?

What This Part Should Cover Guidance

  • Measurable quality checks before and after, and reconciliation of what changed
  • Where the model runs and what data it may see: minimization, masking, synthetic samples
  • Security review of generated code, access control, and auditing of prompts and outputs
  • Compliance obligations for regulated data

What a Strong Answer Covers Guidance

  • Treating the LLM as a fast but fallible assistant whose output is verified, not trusted
  • A concrete prompt structure: role, context, constraints, examples and output format
  • Reviewable, deterministic artifacts (code, SQL, rules) instead of unexplained bulk edits
  • Quality that is measured, with changes that can be reversed
  • Sensitive data kept out of prompts unless the deployment is approved for it

Follow-up Questions Guidance

  • The translated Python passes every test but runs much slower than the Java version. How do you find out why, and would you ask the LLM to fix it?
  • What would it cost in time and tokens to send every row through the LLM, compared with having it write cleanup rules?
  • The model "corrects" a value that was actually right. How does your process catch it?
  • A teammate pasted production customer records into a public chatbot. What do you do now, and what prevents it next time?
Loading comments...