Skip to content

Removing personal information

View as MarkdownOpen in Claude(opens in a new tab)Open in ChatGPT(opens in a new tab)

Vintage’s core data promise is that raw personal information never leaves your browser. Before any file is uploaded, the Remove PII stage asks you two things about each file group: which columns identify the loan and the borrower, and which columns hold other personal information. Identifier columns are replaced by scrambled values in your browser, the other personal-information columns are removed in your browser, and only the cleaned file is sent.

What leaves your browser, and what never does

Section titled “What leaves your browser, and what never does”
WhenWhat
Never leaves your browserAny value from a column you marked as personal information. Any raw Loan ID or Tax ID value. Your original file.
Sent after you confirmA cleaned copy of each file: the columns you kept, under their own headings, with the Loan ID and Tax ID values replaced by protected IDs
Read before you confirmYour column headings (the names only, never the values), so Vintage can suggest which column is the Loan ID, which is the Tax ID, and which look like personal information
Remembered for next timeWhich column names you chose as the Loan ID and Tax ID and which you removed. Never a cell value.

Vintage also checks a sample of each column’s values for patterns that look like email addresses, phone numbers, or Social Security numbers. That check runs inside your browser, and its results only pre-tick boxes on your screen.

Because your original files are never stored, a draft you leave and come back to never reopens at Choose Files or Remove PII. There is no copy of your raw file to return to.

You review this screen once per file group, not once per file. Every file in a group shares one set of column names (see Choosing files), so the decisions you make here apply to every file in the group. When an upload has more than one group, a tab for each type sits above the screen. Confirm and continue confirms the group in front of you and moves to the next group still to be reviewed. When every group is confirmed, the upload begins.

The screen has two parts:

  • Which of your columns have the Loan ID and Tax ID? Two searchable selectors, each listing every column in the file. Vintage pre-selects the columns it believes hold each identifier.
  • Select PII columns to remove before upload. A checkbox for every other column. Columns Vintage believes hold personal information are pre-ticked and tagged Likely PII. Tick or untick any column.

Confirm and continue is available once a Loan ID column is chosen.

The Loan ID is the one identifier that ties a row to a loan. It is how Vintage joins a loan’s snapshot rows, its origination facts, and its transactions together, and how it recognizes the same loan in next month’s file. It is the only field Vintage requires: if no column is chosen, the screen says “Select a Loan ID column to continue.”

The Loan ID column is scrambled, not removed. Its values are replaced in your browser by protected IDs (below), and the scrambled column is uploaded so loans can be joined.

The Tax ID is the borrower’s tax identifier: a Social Security number, a taxpayer identification number (TIN), or an employer identification number (EIN). It is optional. If your file has no such column, choose None.

When you do designate one, it is scrambled the same way as the Loan ID and stored, and it serves as a fallback join key: Loan ID is the primary key that ties rows to a loan, and the Tax ID is the fallback. The Portfolio screen shows whether a Tax ID is on file for each loan, never the value itself.

A protected ID is the scrambled value that replaces a Loan ID or Tax ID. It has two properties that matter:

  • It is one-way. A protected ID cannot be turned back into the identifier it came from. Vintage never holds your real loan numbers or tax identifiers.
  • It is consistent. The same Loan ID always scrambles to the same protected ID, so the loan in this month’s file is recognized as the loan in last month’s file, and a snapshot row joins to its origination row and its transactions.

Everywhere Vintage shows a loan’s identity, it shows the protected ID. You can search the Portfolio screen by protected ID; a raw loan number is never displayed. See The Portfolio screen.

How a Loan ID is normalized before it is scrambled

Section titled “How a Loan ID is normalized before it is scrambled”

Core systems and spreadsheets write the same loan number in different ways: padded with zeros in one export, trimmed in another, with a stray space in a third. If each spelling scrambled differently, one loan’s history would split into several loans. So before a Loan ID is scrambled, Vintage applies three normalizations:

  • All whitespace is removed, including spaces inside the value.
  • Letters match regardless of case.
  • Leading zeros are removed.

Everything else stays significant: separators such as - and / inside an ID, and every digit after the first non-zero character. Genuinely different IDs stay different.

As written in your fileRead as the same loan as
001234512345
12 34512345
ab-1001AB-1001
12345-01not 1234501 (the hyphen is kept)
123450not 12345 (trailing digits are kept)

A Tax ID is normalized by its own rule: every character that is not a letter or a digit is removed, because separators in a tax identifier carry no meaning. 12-3456789 and 123456789 are the same Tax ID.

Some exports fill an empty identifier with a placeholder. Vintage treats these exactly like an empty cell when they appear in a Loan ID or Tax ID column:

  • the blank-like tokens Vintage treats as “unknown” everywhere: N/A, N-A, NA, NULL, NONE, EMPTY, and a lone - (matched regardless of case and spacing);
  • an all-zeros value, such as 0, 000000, or 000-00-0000;
  • for Tax IDs, punctuated forms of the same tokens, such as N.A. or N / A.

A row with a placeholder gets no protected ID. The reason is that scrambling is consistent: if N/A were scrambled, every N/A row in every file would receive the same protected ID, and rows belonging to many different real loans would silently merge into one.

A row with no usable Loan ID is still accepted as part of the upload, but it cannot be linked to a loan. The Review reports how many rows are “without a Loan ID” so you can see them. See Reviewing before you finalize.

Any column you tick under Select PII columns to remove before upload is removed in your browser and never sent: borrower names, addresses, phone numbers, email addresses, dates of birth, and anything else you do not want to leave your machine.

Vintage pre-ticks the columns it believes hold personal information, from two signals: the column headings, and the in-browser check of sample values for email, phone, and Social Security number patterns. Every pre-ticked column is tagged Likely PII, and every one is your decision to keep or change. A column you choose as the Loan ID or Tax ID is not offered for removal, because it is scrambled rather than dropped.

Removing a column is permanent for that upload: its values are not in the uploaded file, so they cannot be mapped or modeled later. Columns you keep but do not want modeled can instead be excluded at Column Mapping, which keeps the data but leaves it out of the measurements. See Mapping columns.

Back to files returns you to Choose Files with your staged files still in place, so you can remove one you added by mistake or add one you forgot. It is never disabled, including on a file the step cannot accept.

Going back discards the choices made on this screen in this session. The set of files may change, so the choices are worked out again rather than carried onto files they were never made for. When you return, the screen fills in again from Vintage’s suggestions and from what it remembers of your previous uploads.

Vintage records your decisions as column names: which named column is the Loan ID, which is the Tax ID, and which named columns to remove, together with the exact set of column names you reviewed. Because every file in a group carries the same names, the decisions apply to each file whatever order its columns are in.

Before any of a file’s bytes leave your browser, Vintage checks that the file’s columns are exactly the set you reviewed. If they are not, no file of that group is cleaned or sent in that pass. Files of other groups finish normally, and you are returned to Remove PII for that group with a message naming the file that did not match. A file added to a group after you reviewed it always needs the group reviewed again; a file is never uploaded under a review made for different columns.

Vintage remembers your choices between uploads

Section titled “Vintage remembers your choices between uploads”

If you upload the same monthly file every month, you should not have to re-select the Loan ID and re-tick the same columns every time. So Vintage remembers, for each file type, which column was the Loan ID, which was the Tax ID, and which columns you removed, and pre-fills them the next time it recognizes the file.

Recognition compares the file’s columns with one previous accepted upload of the same type: the one with the most columns in common, and the most recent when two are tied. It is never a blend of several uploads, which could revive a choice from a format you have since stopped using. A notice at the top of the screen says what happened:

NoticeWhenWhat is pre-filled
Same columns as a previous uploadEvery column matches the earlier file, which the notice namesYour earlier Loan ID, Tax ID, and removed columns
Mostly the same as a previous uploadMost columns match; some are newThe choices for the columns it recognized. The new columns are named, because they are exactly where unreviewed personal information could be.
New file formatYou have uploaded this type before, but none of it matchesNothing. Review every column.
(no notice)Your first upload of this file typeVintage’s suggestions only

Three rules keep this memory safe:

  • “None” is remembered as an answer. If you chose None for the Tax ID, the next recognized file comes back with None rather than a fresh guess.
  • A column you removed is never pre-filled as an identifier. Choosing a column as the Loan ID or Tax ID takes it out of the removal list, so pre-filling it that way would quietly undo your decision to remove it.
  • Memory can only add to the removal list. The automatic check for personal information still runs on every file. If it flags a column you kept last time, the column is pre-ticked. Recognition is a convenience; the check is a safety net, and the convenience cannot switch the safety net off.

Only column names and designations are remembered. A removed column’s contents never leave your browser; Vintage records only that a column with that name was removed, which is what it needs to offer the same choice again.

As a backstop, after a cleaned file arrives Vintage checks a sample of its values for any column that still looks like personal information, such as an email address, phone number, or Social Security number you forgot to remove. A column it is confident about is excluded automatically at Column Mapping, and the screen says so plainly:

Columns Email and Home Phone were dropped during upload. The system detected they may contain personal information. You can re-include them below.

The column carries a note (“Auto-excluded: values look like PII”) and you can re-include it if it is safe. The exclusion and its reason are kept, so a draft reopened later shows the same notice. The protected Loan ID and Tax ID columns are never flagged. When this check cannot run for a file, it is skipped for that file; nothing cruder stands in for it.