Organizing Code and Repositories: Functions, Classes, Comments, GitHub Files
Company: Wells Fargo
Role: Data Scientist
Category: Software Engineering Fundamentals
Difficulty: hard
Interview Round: Technical Screen
When you write code, how do you organize it into functions, classes and files? Do you write comments, and what do they say? When you upload a project to GitHub, which files does the repository contain?
The question came up in an interview for an applied AI research internship, so a research or machine learning codebase is a natural example to use.
### Clarifying Questions
- Is the interviewer asking about research prototypes, production code, or both?
- Will the repository be public, or shared only within a team?
- Does the project involve data, model checkpoints or credentials that should not be committed?
### Part 1 — Functions, classes and files
How do you decide what becomes a function, what becomes a class, and how the code is split across files and folders?
```hint Reasons to change
Group code by what would force it to change, and ask which pieces you would want to test on their own.
```
#### What This Part Should Cover
- Criteria for functions versus classes, with an example of each
- A module and folder layout by responsibility, with one-way dependencies
- Separation of configuration, library code and entry points
- Where notebooks fit
### Part 2 — Comments
Do you write comments? What belongs in a comment, and what does not?
```hint What the code cannot say
Ask what a reader could not work out from the code itself, however clean it is.
```
#### What This Part Should Cover
- Comments that explain intent and decisions rather than mechanics
- Docstrings for public interfaces, including shapes and units in numerical code
- Type hints as checked documentation
- Keeping comments accurate as the code changes
### Part 3 — What goes in the repository
When you upload the project to GitHub, which files are in the repository, and what stays out?
```hint A stranger clones it
Picture someone cloning the repository onto a fresh machine. What do they need to install it, run it and reproduce your results, and what must never reach them?
```
#### What This Part Should Cover
- Documentation, license and dependency specification
- Source code, tests, configuration and automation
- What must stay out (secrets, large data, generated files) and how those are handled instead
- Reproducing the results from the repository alone
### What a Strong Answer Covers
- Concrete, principled organization rules illustrated with a small example
- A clear position on comments that favors intent over narration
- A complete, reproducible and safe repository layout
- Awareness of how research code differs from production code, and when to tighten it
### Follow-up Questions
- How do you keep notebooks from becoming the place where the real logic lives?
- You accidentally pushed an API key. What do you do?
- How do you make an experiment reproducible from the repository alone?
- When is splitting a growing module into a package worthwhile, and when is it premature?
Overview: A software practice question asks how you organize code into functions, classes and files, how and when you write comments, and which files a GitHub repository should and should not contain. It tests modular design judgment, documentation habits, and reproducible, secure project hygiene.
Read the full Wells Fargo Data Scientist interview experience this question came from