Payment / Invoice Reconciliation (problem as I reconstructed it)
Part 1: match by the invoice ID in the memo. Parse the payment and invoice strings, use the invoice ID in the payment's memo to find the matching invoice, and return the string the problem specifies. A payment has a payment ID, an amount and a memo; an invoice record has an invoice ID, an amount and a due date.
Part 2: add amount matching. Keep the Part 1 memo matching, and add a rule that matches an invoice by exact amount. If several invoices have the same amount, pick the one that is due earliest. The output format is the same as Part 1, and the original test cases still have to work.
Part 3: add a "forgiveness" range match. There is now an amount tolerance (forgiveness), so you can look for an invoice within the allowed range; if there are several candidates in range, pick the one due earliest. At the end the interviewer also asked about the priority when both an exact amount match and a range match exist.
Tests and follow-ups from the interviewer:
- "could you invoke the method with some example inputs to make sure it works?"
- "Do you mind just doing one more, like the uh, that bottom example?"
- "you want to also make sure that you're still supporting the original use case as well."
- "We still want those to pass in addition to the new requirements."
- "Yeah, see if you can find any more edge cases."
- "can you think of any other things that you could be missing with this implementation?" / "Any potential bugs?"
- "What if there is an exact match with forgiveness applied?"
- "how would we prioritize that?"
Bug Squash
The task: in the Python requests repo, track down the problem with resending the request body when a POST goes through an HTTP 307 redirect, and make the three tests below pass. After a 307 redirect the request has to stay a POST and send the same body as the first request.
| Input | State before calling POST | What both the first POST and the redirected POST should send |
|---|---|---|
Plain string 'test' | No stream cursor | test |
io.BytesIO(b'test') | Cursor at 0 | test |
io.BytesIO(b'hello world') | read(5) already called, cursor at 5 | world, leading space kept |
You have to find out why the plain string passes while the seekable stream times out after the redirect, and the fix also has to support a stream that the caller has already partly read.
The test code given with the problem:
def test_HTTP_307_ALLOW_REDIRECT_POST_DATA(self, httpbin):
r = requests.post(httpbin('redirect-to'), data='test', params={'url': httpbin('post'), 'status_code': 307}, timeout=5)
assert r.status_code == 200
assert r.history[0].status_code == 307
assert r.history[0].is_redirect
assert r.json()['data'] == 'test'
def test_HTTP_307_ALLOW_REDIRECT_POST_WITH_SEEKABLE(self, httpbin):
byte_str = b'test'
r = requests.post(httpbin('redirect-to'), data=io.BytesIO(byte_str), params={'url': httpbin('post'), 'status_code': 307}, timeout=5)
assert r.status_code == 200
assert r.history[0].status_code == 307
assert r.history[0].is_redirect
assert r.json()['data'] == byte_str.decode('utf-8')
def test_HTTP_307_ALLOW_REDIRECT_POST_WITH_PARTIAL_SEEKABLE(self, httpbin):
data = io.BytesIO(b'hello world')
data.read(len("hello"))
r = requests.post(httpbin('redirect-to'), data=data, params={'url': httpbin('post'), 'status_code': 307}, timeout=5)
assert r.status_code == 200
assert r.history[0].status_code == 307
assert r.history[0].is_redirect
assert r.json()['data'] == ' world'
What the bug is: the first POST reads the request body stream to the end, and when the 307 redirect resends the request it still uses the same stream object, but nobody put the cursor back to where it was before the first send. read() does not delete the data in the buffer, it only moves the cursor. After the first send of BytesIO(b'test') the cursor goes from 0 to 4; reading again from the current position gives only b''. The redirected request still carries Content-Length: 4 but cannot supply those four bytes, so the server waits for the body, the client waits for the response, and what you finally see is a ReadTimeout.
A plain string has no cursor that gets advanced by reading, so it can be sent again without trouble. The partly read stream is the better test of whether the fix is right: BytesIO(b'hello world') had five bytes read before the call, so the first send was the six bytes of world, which means the resend has to go back to position 5, not always to 0.
Full stream:
before send pos=0 -> first send b'test' -> pos=4
original resend logic: still reads from pos=4 -> b'', but Content-Length is still 4
required: restore pos=0, then send b'test'
Partly read stream:
before send pos=5 -> first send b' world' -> pos=11
original resend logic: still reads from pos=11 -> b'', but Content-Length is still 6
required: restore pos=5, then send b' world'
Where the bug is: it involves request preparation and copying in models.py, and the redirect resend in sessions.py. Below are the defects in the original version of the problem and the relevant code.
The initial position of the stream is never saved. In PreparedRequest.prepare_body in models.py, the stream branch just uses body = data and sets Content-Length from the remaining length. The original version does not record the cursor before the first send, so afterwards nobody knows where to replay from.
Copying the request does not copy the stream or reset the cursor. In PreparedRequest.copy in models.py, p.body = self.body makes the two requests share the same stream object. That assignment itself does not need to become a deep copy; the issue is that the resend logic cannot assume the shared stream is still at the start.
The place that is actually missing the cursor restore is just before the resend. SessionRedirectMixin.resolve_redirects in sessions.py keeps the body and the related headers for 307/308 and then calls self.send(...) again. The original version does nothing in between to restore the read position of the body.
The key code from the original version:
# models.py - PreparedRequest.prepare_body
if is_stream:
body = data
# the original version does not save the initial stream position here
# models.py - PreparedRequest.copy
p.body = self.body
# sessions.py - SessionRedirectMixin.resolve_redirects
if resp.status_code not in (codes.temporary_redirect, codes.permanent_redirect):
purged_headers = ('Content-Length', 'Content-Type', 'Transfer-Encoding')
for header in purged_headers:
prepared_request.headers.pop(header, None)
prepared_request.body = None
# 307/308 keep the body; the original version does not rewind before calling self.send(...) again
These are excerpts from the original version, and the comments are my own explanatory additions, not a full function.
Fix idea (pseudocode). The key is to record and restore the position from before the first send, not to always seek(0):
prepare the request:
if body is a replayable stream:
request.body_start = body.tell()
copy the request:
copied.body = request.body
copied.body_start = request.body_start
handle a redirect that keeps the body:
if body is a replayable stream:
body.seek(copied.body_start)
send copied again
Integration - BikeMap
Read GeoJSON data of a bike ride and generate a map image through the staticmap API. The problem has three parts, and each later part extends the code from the earlier one.
Input: ride-simple.json, the bike ride track data.
API: POST to the staticmap endpoint, map.png.
Request body: JSON with the map parameters, plus the markers and paths added later.
Response body: the binary data of a PNG image, not JSON.
Output: Part 2 saves bikemap2.png, Part 3 saves bikemap3.png, and you open them manually to look at them.
Part 1: Parse JSON (problem as I reconstructed it). Read ride-simple.json, pull the first ten coordinates out of the bicycle trip data and print them.
The top level of the file is a FeatureCollection. The track is in the geometry of features[0], with type LineString, and the coordinate array is at features[0]["geometry"]["coordinates"]. GeoJSON coordinate order is [longitude, latitude], longitude first and latitude second. staticmap uses {"lat": latitude, "lon": longitude}, so do not swap them when building the request.
The local input file has 495 coordinates, i.e. 494 adjacent segments. The first coordinate is [-122.2851, 47.55126], which in staticmap form is {"lat": 47.55126, "lon": -122.2851}. These are notes on the structure of the input file; the full English original of Part 1 was not saved.
Part 2, the original problem text:
Write a program that sends an HTTP POST request with a
JSON request body to staticmap.
Save the binary response image data to bikemap2.png.
Manually open this image in a web browser or an image previewer, using the File menu or the command line.
Write your code in the same file as Part 1.
Use the following map parameters in the POST request data:
{
"center": {
"lat": 47.579,
"lon": -122.31
},
"width": 400,
"height": 600,
"zoom": 13
}
These are the same center, width, height, zoom from staticmap_example.json.
Part 3, the original problem text:
Let's draw a map of GeoJSON data using staticmap.
You should:
Write code to render markers and paths on a map.
Use the coordinates in ride-simple.json.
Plot the following Markers on the map:
The first coordinate: a "white" marker labeled Start.
The last coordinate: a "blue" marker labeled Stripe.
Plot a series of contiguous Paths on the map:
Draw a "blue" path through the first ten line segments between the coordinates.
Draw a "brown" path through the second ten line segments between the coordinates.
Continue alternating blue and brown until
you've used all the coordinates.
Use the same map parameters
(center, width, height, zoom from staticmap_example.json) as before,
save the response image data to bikemap3.png,
and open the image manually.
Extend your previous solution.
Request structure and edge cases for Part 3. On top of the Part 2 center, width, height and zoom, two fields are added:
| Field | Structure and requirements |
|---|---|
markers | An array; each marker has color, label and coord, where coord is {"lat": ..., "lon": ...} |
paths | An array; each path has color and positions, where positions is an array of coordinate objects in track order |
Start: the first coordinate, a white marker labeled Start. End: the last coordinate, a blue marker labeled Stripe.
The color changes every ten segments, not every ten coordinates. Ten contiguous segments need eleven coordinates. The first group is blue, the second is brown, and so on alternating; neighboring groups share the boundary coordinate so the track stays continuous.
The leftover part at the end with fewer than ten segments must be kept too. In this problem 494 segments make 49 full groups, with four segments left over.
The main point for using Python requests, the JSON payload goes through json=data. Here data is a Python dict and the API wants a JSON body, so you write requests.post(url, json=data), not requests.post(url, data=data). The package is called requests.
import requests
url = "<staticmap endpoint>"
data = {
"center": {"lat": 47.579, "lon": -122.31},
"width": 400,
"height": 600,
"zoom": 13,
}
response = requests.post(url, json=data)
response.raise_for_status()
with open("bikemap2.png", "wb") as image_file:
image_file.write(response.content)
This is an example of the request and saving the image, not the full solution that draws the track. Once Part 3 adds markers and paths it still uses json=data, and the output file name changes to bikemap3.png.
| How you write it | What actually happens |
|---|---|
requests.post(url, json=data) | Serializes the dict to JSON; if you do not override it, it sets Content-Type: application/json automatically, which is right for this problem |
requests.post(url, data=data) | The dict is form-encoded, usually as application/x-www-form-urlencoded; it does not become JSON just because the variable is called data or has nested dicts inside |
requests.post(url, data=json.dumps(data), headers={"Content-Type": "application/json"}) | Serializing by hand and setting the type also works, but using json=data directly is simpler for this problem |
Also watch out for these. Pass the dict straight to json=data; do not write json=json.dumps(data), or the already serialized string gets encoded as JSON a second time. Do not pass data= and json= at the same time; a non-empty data makes Requests ignore the json argument. A JSON request does not mean the response is JSON too. Here the response is a PNG, so use response.content, do not call response.json(), and do not save the image through response.text. Open the file with "wb", and call raise_for_status() first so that an error response like a 400 does not get saved as a .png. If you hit a 400, look at the status code and the error body first, then check whether the request body is really JSON; just adding a JSON header will not turn the form content of data=data into JSON.
Excerpt of the track handling code, which covers the coordinate conversion, the markers and the path grouping. coordinates is the coordinate array from the input file and data is the map parameter dict from above; this is the key fragment, not what was submitted in the interview.
# GeoJSON: [longitude, latitude] -> staticmap: {lat, lon}
positions = [{"lat": lat, "lon": lon} for lon, lat in coordinates]
data["markers"] = [
{"color": "white", "label": "Start", "coord": positions[0]},
{"color": "blue", "label": "Stripe", "coord": positions[-1]},
]
data["paths"] = [
{
"color": "blue" if (i // 10) % 2 == 0 else "brown",
# ten segments need eleven points, neighboring groups share an endpoint
"positions": positions[i:i + 11],
}
for i in range(0, len(positions) - 1, 10)
]
# the JSON request body uses json=data; the response is image bytes
response = requests.post(url, json=data)
response.raise_for_status()
with open("bikemap3.png", "wb") as image_file:
image_file.write(response.content)
AI Programming Exercise (problem as I reconstructed it)
Using the AI built into the platform, implement transaction risk rule evaluation in Python: given a dict of fields for one transaction and an ordered list of rules, decide accept / block.
Part 1: string conditions and rule order. Support string equality conditions and the two outcomes accept / block. Apply the rules in the order of the rules list, and the first rule that matches decides the result. If no rule matches, the default is accept. The examples allow the field and the string constant to be swapped on either side of the equals sign. The input is guaranteed to be well formed.
Part 2: Boolean and AND / OR. Keeping the Part 1 behavior, add Boolean field conditions and AND / OR. The interviewer also raised an edge-case test: what does the program do if the merchant name itself contains "and" or "or"?
Part 3: more complex combined conditions. Keep extending the condition expressions, involving combinations of parentheses, equality checks and AND / OR.
Discussion
Loading comments…