Skip to content

Assignment 4: Static Analysis and Continuous Integration

Possible Points Due Date
50 pts Friday, October 9th - 11:59pm

Overview

Every test in the repository you are about to receive passes. The app still has six bugs, and one of them hands a member's private email address to anyone who knows how to ask. None of the tests can see them, because tests only check the inputs somebody thought to write down.

In this assignment you add analysis tools to the CI pipeline one at a time. Each one is from a different family in the static and dynamic analysis lecture, and each one finds something the others cannot.

After this assignment you should be able to:

  • Add a static analysis tool to a project and to its CI pipeline, as a check that fails the build
  • Triage a warning: decide whether it is a real bug, a style issue, or a false positive, and act accordingly
  • Explain what a linter, a type checker, and a pattern-based security scanner can and cannot find
  • Use coverage as dynamic analysis, reading what it says was run without mistaking that for what was checked
  • Suppress a false positive the responsible way, with a narrow, justified exception rather than by switching the check off

You will work on Bookshelf, a small lending-library API written in Python with Flask. Assignment 3 was about the pipeline carrying code to users; this one is about what that pipeline checks along the way.

Using AI Coding Assistants

You are welcome to use AI assistants on this assignment. If you do, you must record it: add an entry to your Collaboration Log noting which tool you used and what for. Review the full Policy on the Use of AI Coding Tools before you start. You are responsible for anything an AI tool produces, so read and verify it. Note that this assignment grades your triage: an assistant that silences a warning instead of fixing its cause will cost you points, not save you time.

Prerequisites

  • uv, the Python project manager (uv --version to check). It installs the right Python for you, so you do not need to install Python separately
  • A GitHub account, already authenticated for git on your machine
  • The Git workflow from Assignment 1: branch, commit, pull request, merge
  • How a CI workflow runs on a pull request, from Assignment 3

Windows

uv runs natively on Windows, and everything in this assignment works in PowerShell. If you already use WSL for the course, use it here too; just install uv inside WSL rather than on the Windows side.

No Python experience?

You do not need much. The app is about 150 lines, every function has type annotations and a docstring, and the bugs are the kind you would recognize in any language. If a construct is unfamiliar, the official Python tutorial covers everything used here.

Getting your repository

Your repository comes from UCF Code Classroom, the same app you used for Assignments 1 through 3. You are already enrolled, so there is no new link to accept: sign in, open Assignment 4, and your repository will be there.

This is an individual assignment.

If your repository is not there

Post on Ed Discussions. Do not create your own repository or copy Bookshelf somewhere else - Code Classroom collects the repository it created for you, and work anywhere else will not be picked up.

Once you have it, get it running:

git clone https://github.com/UCF-CEN-5016/<your-assignment-4-repo>.git
cd <your-assignment-4-repo>
uv sync
uv run pytest

uv sync creates a .venv directory containing exactly the versions pinned in uv.lock, downloading Python 3.12 first if you do not have it. uv run runs a command inside that environment, so there is no virtualenv to activate. CI runs uv sync --locked, which refuses to proceed if the lockfile is out of date, so your machine and CI install the same thing.

Then start the app and try it:

uv run flask --app bookshelf.app run --port 8000
curl http://127.0.0.1:8000/api/books/1
curl "http://127.0.0.1:8000/api/books/search?q=code"
curl "http://127.0.0.1:8000/api/loans/1/fee?returned=2026-10-06"

The README.md in the repository lists every endpoint and explains how the code is laid out. Read it now.

Why port 8000?

Flask's default is 5000, but on macOS the AirPlay Receiver already listens there and answers every request with a 403 Forbidden, which looks exactly like a bug in your app. Using 8000 sidesteps it on every platform.

How you will work

Parts 2 through 5 each add one tool. Do them in order, and do each one as its own pull request:

  1. Branch off an up-to-date main.
  2. First commit: add the tool, its configuration, and its CI job - and nothing else. Push, and open the pull request. CI goes red, and the job's log lists what the tool found. That red run is part of your grade: it is the evidence that the check actually catches something.
  3. Further commits: triage and fix the findings until CI goes green.
  4. Merge, then pull main before starting the next part.

Do not force-push over that first commit, and do not start the next part on a branch cut before the previous part merged. Each part changes pyproject.toml and uv.lock, so parallel branches will conflict.

Required checks

On a team project you would make each of these jobs a required status check in branch protection, so GitHub refuses to merge a pull request until it is green. Code Classroom manages your repository's settings, so you cannot turn that on here and you do not need to. Hold yourself to the rule anyway: merge only on green.

Triage is the assignment

A tool's output is a list of claims, not a list of instructions. For every finding, decide which it is:

  • A real bug. Fix the code, and, where you can, prove the bug with a test that fails before your fix.
  • A style or maintainability issue. Fix it, usually with the tool's own autofix.
  • A false positive. The tool is wrong about this line. Suppress it on that line only, with a comment explaining why the tool is wrong. Never disable the rule for the whole project to make one warning go away.

The lecture's warning applies throughout: every static analysis is incomplete, unsound, or both (Rice's theorem), so "be prepared for MAYBE". The judgment is yours, and it is what you are graded on.

Silencing is not fixing

Making a warning disappear without addressing its cause - a bare # type: ignore, a cast(), # noqa or # nosec on a real bug, deleting or editing a test so it stops failing, or removing a rule from the configuration - gets no credit for that finding.

Part 1: Run it and read it (5 pts)

Before any tool touches the code, look at it yourself. It is small enough to read in one sitting.

  1. Run the tests and start the app as above.
  2. Open .github/workflows/ci.yml. There is one job, test. You will add a job per part.
  3. Spend fifteen minutes hunting for bugs by reading the code and trying the API with curl. Write down anything you find, or suspect, in a new file, REFLECTION.md, under a heading ## Part 1: By hand. It does not matter how many you find; what matters is that you record it honestly before the tools tell you, because the reflection compares the two.

Commit REFLECTION.md to main directly or through a pull request, your choice.

Deliverable: a screenshot of uv run pytest passing locally, and your Part 1 notes in REFLECTION.md.

Part 2: Linting and formatting with Ruff (10 pts)

Ruff is a linter and a formatter in one tool. In the lecture's terms it covers the first two families: linting (shallow syntax checks and style) and pattern-based bug detection. By default it enables only a conservative set of rules, so you will turn on two more families: B (bugbear, patterns that are likely bugs) and I (import ordering).

Add it as a development dependency. It is a tool you use while developing, not something the app needs to run:

git switch -c add-ruff
uv add --dev ruff

Then add its configuration to the end of pyproject.toml:

[tool.ruff.lint]
select = ["E", "F", "B", "I"]  # pycodestyle, Pyflakes, bugbear, isort

And add a lint job to .github/workflows/ci.yml, alongside test:

  lint:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7
      - uses: astral-sh/setup-uv@v10.2.0 #(1)!

      - name: Install dependencies
        run: uv sync --locked #(2)!

      - name: Lint
        run: uv run ruff check --output-format=github . #(3)!

      - name: Check formatting
        run: uv run ruff format --check . #(4)!
  1. Installs uv on the runner. uv then reads .python-version and installs that Python itself, so there is no separate setup-python step.
  2. --locked makes this fail if uv.lock does not match pyproject.toml, which is how you find out you forgot to commit the lockfile.
  3. --output-format=github prints each finding in the format GitHub Actions understands, so it appears as an annotation on the exact line in the pull request's Files changed tab.
  4. --check reports files the formatter would change without changing them. CI should never rewrite your code; it should tell you, and you run the formatter yourself.

Commit those three changes, push, and open the pull request. Then work through the findings locally:

uv run ruff check .          # list the findings
uv run ruff check --fix .    # apply the fixes Ruff is sure are safe
uv run ruff format .         # reformat every file

Ruff reports four findings. Most are style. One is a real bug, and it is visible from the API: request two different books, one after the other, and compare their tags. Before you fix it, work out why it happens - the rule's documentation page explains it - and add a test to tests/test_routes.py that fails because of it.

Read the rule, not just the message

Every Ruff finding has a code like F401 or B006. Search it on the rules page to see why the rule exists, an example, and the recommended fix. That page is how you tell a style nit from a bug.

The formatter changed files I did not touch

That is expected, and it is why you run it once over the whole codebase rather than a file at a time. From then on, the Check formatting step keeps it that way.

Deliverables: the merged pull request, showing a failing lint run followed by a passing one, and a test that demonstrates the real bug.

Part 3: Type checking with mypy (10 pts)

Python does not check types before it runs, but Bookshelf is fully annotated, and mypy can use those annotations to check, without running anything, that every value is used consistently with its type. This is the lecture's type-annotation validator family: a conservative analysis that can prove certain errors absent, as long as the annotations are right.

git switch main && git pull
git switch -c add-mypy
uv add --dev mypy
[tool.mypy]
strict = true              # every check mypy has, including missing annotations
files = ["bookshelf"]      # the application code; the tests are not annotated

This time, write the CI job yourself. Call it types. It is the same four setup lines as lint, followed by one step that runs uv run mypy. Commit, push, and open the pull request.

mypy reports six errors from two root causes. Both are real bugs, and both turn into a 500 Internal Server Error for a request the tests never make. For each one, find that request and try it with curl before you fix anything.

When you fix them:

  • A missing record is not a server error. Answer as GET /api/loans/<id>/fee already does for an unknown loan: a 404 with a JSON error.
  • Query parameters always arrive as strings. Flask can convert one for you; see the type argument of request.args.get.
  • Add a test for each bug to tests/test_routes.py.

Six errors, two bugs

The five union-attr errors all come from one line. mypy reports every place a possibly-None value is used, not the one place it could have been checked. When a tool reports a cluster of findings together, look for the single cause before fixing each symptom.

Deliverables: the merged pull request, showing a failing types run followed by a passing one, and a test for each of the two bugs.

Part 4: Security scanning with Bandit (10 pts)

Bandit is a pattern-based security scanner: it matches the code against a catalogue of known-dangerous patterns. It does not track data through the program the way the lecture's taint analysis does, so it cannot tell whether dangerous input actually reaches a dangerous call. Expect it to be right sometimes, wrong sometimes, and unsure often - its reports include a confidence for exactly that reason.

git switch main && git pull
git switch -c add-bandit
uv add --dev bandit

Write a CI job called security that runs uv run bandit -r bookshelf. Commit, push, and open the pull request.

Bandit reports three findings. Two are real and one is a false positive. Deciding which is which is the point of this part.

  1. The SQL injection. This is the bug from the overview. Before you fix it, write a test in tests/test_routes.py that exploits it: a search that makes the API return a member's email address. The members table and its columns are in bookshelf/db.py, and the query you are attacking is right in front of you, so this is a puzzle about SQL, not about hacking. The test should send your exploit and assert that no email address comes back. Run it against the vulnerable code and watch it fail: that failure is your proof that the warning is real. Then fix the query using a parameterized query, never by escaping quotes yourself.

    A second, innocent witness

    You do not have to be an attacker to trip this bug. Search the catalogue for a book whose title contains an apostrophe.

  2. The other real finding. Read Bandit's explanation, and decide how the app should behave instead. Keep the ability to turn the behaviour on locally when you want it, from an environment variable.

  3. The false positive. Explain, in a comment on the line above it, why the warning does not apply here, then suppress that one rule on that one line with # nosec <rule id>. Put nothing after the rule id: Bandit reads any further words as more rule ids and warns about them.

Deliverables: the merged pull request, showing a failing security run followed by a passing one, and your injection test.

Part 5: Coverage as a gate (10 pts)

Everything so far has been static: the tools read the code without running it. Coverage is dynamic analysis. It instruments the program, runs your tests, and records which lines and branches actually executed. It never over-approximates - what it reports definitely happened - but it can only tell you about the runs your tests perform.

git switch main && git pull
git switch -c add-coverage
uv add --dev pytest-cov
[tool.coverage.run]
branch = true              # measure both outcomes of every if, not just lines
source = ["bookshelf"]

[tool.coverage.report]
show_missing = true        # list the lines and branches that never ran
fail_under = 100           # and fail the run if there are any
exclude_also = ['if __name__ == "__main__":']

Change the Run tests step of the existing test job to uv run pytest --cov, commit, push, and open the pull request. The test job now fails even though every test passes, because the suite never exercises every branch.

Read the Missing column. Entries like 36 are lines that never ran; entries like 62->60 are branches where one outcome never happened. For each one, write a test that exercises it, starting from what the code is supposed to do.

At least one of those tests will fail. That is the last bug. Before you touch the code, re-read the docstring of late_fee_cents(): it is the specification, and when the code and the docstring disagree, the code is wrong. Fix the code, not the test.

Why 100%, and why not everywhere

100% branch coverage is a reasonable bar for a 150-line app, and it makes the gate unambiguous. In a large codebase it is usually a poor target: the last few percent tend to be defensive code that is expensive to reach, and a number that high invites tests written to touch lines rather than to check them. Notice, too, what the gate does not guarantee. A test that runs the fee cap and asserts nothing would satisfy it. Coverage tells you what ran, not what was checked.

Deliverables: the merged pull request, showing a failing test run followed by a passing one, and the tests you added.

Part 6: Reflection (5 pts)

Add a ## Part 6: Reflection section to REFLECTION.md, 300-500 words, answering:

  • Which tool found which bug? List all six bugs you fixed, each with the tool that reported it and one sentence on what a user would have seen. Then compare that list with your Part 1 notes. What did reading find that no tool did, if anything, and what did the tools find that you missed?
  • Why could no single tool have found them all? Pick one bug that only a static tool found and one that only a dynamic one found, and explain why the other kind of analysis missed it.
  • The false positive. Why was Bandit wrong, and why would turning the rule off for the whole project have been the wrong fix?
  • What is still missing? Describe one kind of bug Bookshelf could still have that all four of these checks would let through, and say what would catch it.

Deliverable: REFLECTION.md committed to main.

Extra Credit (5 pts)

CI finds problems minutes after you push. The lecture's point that analysis should be fast and continuous suggests catching them before you commit at all.

Set up pre-commit so that git commit runs Ruff's linter and formatter on the files you are committing, and refuses the commit if they fail. Use the official Ruff hooks. You do not need to install pre-commit into the project: uvx pre-commit install runs it without adding it as a dependency.

Commit .pre-commit-config.yaml, add a section to README.md telling the next developer how to enable it, and add a screenshot to your submission of a commit being refused.

Turning in the Assignment & Grading Rubric

Submit on Webcourses:

  • The URL of your GitHub repository
  • Screenshot: uv run pytest passing locally (Part 1)
  • Screenshot: a commit refused by pre-commit (Extra Credit, if attempted)
  • Your AI collaboration log, if you used AI tools

Everything else is read from your repository. Before the deadline, make sure that all four pull requests are merged into main, each still showing its failing run followed by its passing one, and that REFLECTION.md is committed. main should end with all four jobs - test, lint, types, and security - passing.

The grading for this assignment breaks down as follows:

  • Part 1: Run It and Read It - (5 points)
  • Part 2: Linting and Formatting with Ruff - (10 points)
  • Part 3: Type Checking with mypy - (10 points)
  • Part 4: Security Scanning with Bandit - (10 points)
  • Part 5: Coverage as a Gate - (10 points)
  • Part 6: Reflection - (5 points)
  • Extra Credit: Pre-commit Hooks - (5 points)

Within Parts 2 through 5, points are for: a CI job that fails on the original code and passes on yours; each finding correctly triaged and fixed at its cause; and the tests that each part asks for.