From Files to Features: What a Machine Learning Model Actually Sees

A JPEG, a song, a sentence and a social network reach a model as numbers. See the four layers between file and model input, and why that choice shapes results.

Sylla N'falyOct 01, 202618 min read

What does a machine learning model actually see when you give it a photograph, a song, a video, a spreadsheet or a social network?

Not the photograph. A JPEG on disk is a few hundred kilobytes of compressed bytes, organised for storage, not for arithmetic. Before any model touches it, a decoder has rebuilt a grid of pixel values, and some preprocessing code has probably resized it, reordered its axes and rescaled its numbers. The song, the sentence and the social network go through the same kind of journey, with different tools and very different results.

A popular shortcut summarises the journey as "everything becomes a matrix". It contains a real insight and a real error. The insight: every machine learning system computes on numbers, so every kind of data must be turned into numbers first. The error: those numbers do not all take the same shape, and choosing their shape is often the most consequential decision in the whole project.

The more accurate version of the idea is less catchy and more useful. What is universal is not the matrix. It is the need for a numerical representation suited to the computation you want to perform.


A file is not what the model sees

Four layers sit between the world and a model. They are often blurred together, and most data bugs live in the gap between two of them.

  1. Storage. How information is kept or transmitted: CSV, Parquet, JSON, JPEG, PNG, WAV, MP4, rows in a database. A storage format is optimised for size, durability, streaming or interoperability. It is not optimised for learning.
  2. Decoding. How software turns stored bytes back into usable values: a CSV parser, an image codec, an audio decoder, a database driver. Decoding can lose information (a lossy JPEG never gives back the original scene) and it can make silent choices (which channel comes first, which sample rate to use).
  3. In-memory numerical representation. What exists in RAM after decoding: arrays, matrices, tensors, sequences of integers, sparse structures, graph objects. This layer is shaped by libraries (NumPy, pandas, PyTorch), not by the model.
  4. Model representation. What a specific architecture expects: normalised tensors in a given axis order, token IDs, patches, embeddings, node and edge feature tables. Two models can need different representations of the same image.

The gap between layers is easy to measure. A full HD photograph, decoded to 8-bit RGB, occupies 1920 × 1080 × 3 = 6,220,800 bytes in memory, whatever its size on disk, because decoding undoes the compression. Storage is about keeping information cheaply. The in-memory layer is about making it addressable. The model layer is about making it learnable.

A small vocabulary note before going further. In NumPy and PyTorch, a tensor is simply a multidimensional array of numbers with a shape and a data type: a vector has one axis, a matrix two, an RGB image three, a batch of images four. That practical meaning is the one used here. It does not imply that every system stores all its data as one dense tensor; some of the most important representations below are sparse lists or irregular structures.


Six kinds of data, the same seven questions

The rest of the article walks through six families of data with the same questions: what exists in the world, how it is stored, how it is decoded, what it looks like in memory, which transformations are common, what a model actually consumes, and what problem this enables. The repetition is deliberate. Once you can answer these seven questions for a new data type, you understand its pipeline.

Every printed output below was produced by running the code shown (Python 3.11, NumPy 2.4, pandas 3.0, Pillow 12.3, OpenCV 5.0, librosa 0.11, PyAV 18.1, tokenizers 0.23, PyTorch 2.14, NetworkX 3.6). The test files are small ones generated for the occasion, so the shapes are easy to check by hand.

Tables and time series

QuestionAnswer
In the worldmeasurements, transactions, events, each with a time and several attributes
Stored asCSV, Parquet, database tables, logs
Decoded bya CSV parser, a Parquet reader, a database driver
In memorya DataFrame (one type per column), then a numeric array X∈RT×FX \in \mathbb{R}^{T \times F}
Transformed bytype conversion, encoding of categories, scaling, sliding windows
Model consumesfeature vectors (one row per example) or windows of shape (N,T,F)(N, T, F)
Enablespredictive maintenance, demand forecasting, energy load forecasting

Here TT is the number of time steps, FF the number of features measured at each step and NN the number of training examples.

Tabular data looks like it is already numerical. It rarely is. Take six readings from an industrial press:

import io
import numpy as np
import pandas as pd
 
csv_text = """timestamp,machine,temperature_c,vibration_mm_s,status
2026-03-01 08:00,press-01,61.2,2.31,ok
2026-03-01 08:01,press-01,61.9,2.40,ok
2026-03-01 08:02,press-01,63.4,2.95,ok
2026-03-01 08:03,press-01,66.8,4.12,warning
2026-03-01 08:04,press-01,70.1,5.87,warning
2026-03-01 08:05,press-01,69.5,5.20,warning
"""
df = pd.read_csv(io.StringIO(csv_text), parse_dates=["timestamp"])
print(df.dtypes.to_string())
print("whole frame ->", df.to_numpy().dtype)
 
X = df[["temperature_c", "vibration_mm_s"]].to_numpy()
print("X", X.shape, X.dtype)
 
windows = np.lib.stride_tricks.sliding_window_view(X, window_shape=3, axis=0)
windows = windows.transpose(0, 2, 1)
print("windows (N, T, F)", windows.shape)
timestamp         datetime64[us]
machine                      str
temperature_c            float64
vibration_mm_s           float64
status                       str
whole frame -> object
X (6, 2) float64
windows (N, T, F) (4, 3, 2)

The CSV file was text. The parser turned it into five typed columns: a timestamp, two floating-point measurements and two strings. Converting the whole frame to an array gives object, an array of Python objects that no numerical library can multiply efficiently. Selecting the numeric columns gives a clean 6×26 \times 2 matrix of floats.

The last step is where modelling begins. A forecasting model does not learn from one reading; it learns from a short history. Sliding a window of three steps over six readings gives four overlapping examples, each a 3×23 \times 2 slice. Nothing was added to the data. The same numbers were reorganised so that "the last three minutes" became an explicit object a model can consume.

Time is the most important column and the easiest one to waste. A timestamp is not a feature until you decide what it means: an ordering, a gap since the last event, an hour of the day, a day of the week. Each choice tells the model something different about the world.

Images

QuestionAnswer
In the worldlight hitting a sensor
Stored asJPEG (lossy), PNG (lossless), TIFF, DICOM for medical images
Decoded byan image codec, through Pillow, OpenCV or torchvision
In memoryan array H×W×CH \times W \times C, usually 8-bit integers from 0 to 255
Transformed bychannel reordering, resizing, cropping, scaling to [0,1][0, 1], normalisation
Model consumesa batch N×C×H×WN \times C \times H \times W for many CNNs, a sequence of patches for vision transformers
Enablesdefect detection on production lines, medical image analysis, perception for robots and vehicles

HH and WW are the height and width in pixels, CC the number of channels and NN the number of images in a batch.

The usual summary, "an image is H×W×3H \times W \times 3", is true for one common case. A grayscale image has one channel. A PNG with transparency has four. And two popular libraries disagree about the order of the colours:

import numpy as np
import cv2
from PIL import Image
 
img = Image.open("logo.png")              # a 6x4 red PNG, half transparent
a = np.asarray(img)
print("Pillow mode", img.mode, "size (W, H)", img.size)
print("Pillow array", a.shape, a.dtype, "first pixel", a[0, 0].tolist())
 
b = cv2.imread("logo.png")                # default flag
print("OpenCV default", b.shape, "first pixel", b[0, 0].tolist())
Pillow mode RGBA size (W, H) (6, 4)
Pillow array (4, 6, 4) uint8 first pixel [255, 0, 0, 128]
OpenCV default (4, 6, 3) first pixel [0, 0, 255]

Same file, two decodings. Pillow keeps the four channels in RGBA order: full red, no green, no blue, half opacity. OpenCV's default reading drops the alpha channel and stores the colours as BGR, so the red pixel appears as [0, 0, 255]. Note also that Pillow reports size as (width, height) while the array shape is (height, width, channels). Feed OpenCV output to a model trained on RGB images and nothing crashes: the model simply sees a world where red and blue are swapped.

PyTorch convolutional models usually expect yet another layout: channels first, scaled to floating point.

import torch
 
rgb = np.array(img.convert("RGB"))                   # (4, 6, 3) uint8
t = torch.from_numpy(rgb).permute(2, 0, 1).float().div(255)
print("torch CHW", tuple(t.shape), t.dtype, "| batch NCHW", tuple(t.unsqueeze(0).shape))
torch CHW (3, 4, 6) torch.float32 | batch NCHW (1, 3, 4, 6)

The decoded array was uint8 values from 0 to 255 in height-width-channel order. The model input is float32 values between 0 and 1, channels first, with a leading batch axis. Most pipelines then subtract a per-channel mean and divide by a standard deviation computed on the training data. None of this is cosmetic: a model trained on normalised inputs produces nonsense on raw ones.

Layout also exists below the shape. PyTorch can store the same (8,3,224,224)(8, 3, 224, 224) batch with channels interleaved in memory (channels_last): the shape stays identical, the strides change, and some convolution kernels run faster. And vision transformers do not consume a pixel grid at all; they cut the image into fixed-size patches and treat them as a sequence of tokens (Dosovitskiy et al., 2020).

Audio

QuestionAnswer
In the worldvariations of air pressure over time
Stored asWAV (usually uncompressed PCM), FLAC (lossless), MP3 or AAC (lossy)
Decoded byan audio decoder, through soundfile, librosa or torchaudio
In memorya waveform: a 1D array of samples x∈Rnx \in \mathbb{R}^{n}, one row per channel if stereo
Transformed byresampling, mixing to mono, short-time Fourier transform, Mel scaling, log compression
Model consumesraw waveform samples, a spectrogram, a log-Mel spectrogram, or hand-made features
Enablesspeech recognition, acoustic event detection, environmental and wildlife monitoring

A microphone measures pressure thousands of times per second. At 44,100 samples per second, two seconds of mono audio is an array of 88,200 numbers. That is the waveform, and nn above is simply duration times sample rate.

The first trap is in the loader:

import librosa
 
y, sr = librosa.load("tone.wav")           # 2 s, 44.1 kHz, 16-bit PCM on disk
print("default", y.shape, y.dtype, sr)
y_native, sr_native = librosa.load("tone.wav", sr=None)
print("sr=None", y_native.shape, sr_native)
default (44100,) float32 22050
sr=None (88200,) 44100

The file contains 16-bit integers at 44.1 kHz. By default, librosa resamples to 22,050 Hz and returns 32-bit floats, so half the samples disappear before you have written a line of modelling code. That default is reasonable for music analysis and documented (librosa), but a model trained at one sample rate and served at another is a classic production bug.

From the waveform, many representations are possible. A short-time Fourier transform cuts the signal into overlapping windows and measures the energy of each frequency in each window:

S = librosa.stft(y)                                    # n_fft=2048, hop_length=512
M = librosa.feature.melspectrogram(y=y, sr=sr, n_mels=128)
print("stft", S.shape, S.dtype, "| mel", M.shape)
stft (1025, 87) complex64 | mel (128, 87)

The STFT has 1+2048/2=10251 + 2048/2 = 1025 frequency bins and 87 time frames (one every 512 samples, plus one because the signal is padded at both ends). Its values are complex numbers: magnitude and phase. The Mel spectrogram keeps only energy and regroups frequencies into 128 bands spaced the way human hearing distinguishes pitch. It is a 128×87128 \times 87 matrix: an image-like object, which is why convolutional networks work well on it.

It is tempting to conclude that "audio becomes a spectrogram". Many systems do exactly that. Whisper, for instance, resamples audio to 16 kHz and computes an 80-channel log-magnitude Mel spectrogram on 25 ms windows with a 10 ms stride (Radford et al., 2022). Others skip the spectrogram entirely: wav2vec 2.0 learns its own representation directly from the raw waveform (Baevski et al., 2020). A spectrogram is a choice. It discards phase, fixes a trade-off between time and frequency resolution, and builds in assumptions about hearing. Sometimes those assumptions help; sometimes a learned front end does better.

Video

QuestionAnswer
In the worlda scene changing over time, often with sound
Stored asa container (MP4, MKV, WebM) holding compressed video and audio streams (H.264, H.265, AV1…)
Decoded bya video decoder, through PyAV, OpenCV, torchvision or Decord, often on dedicated hardware
In memoryframes, one at a time: each an H×W×CH \times W \times C array
Transformed byframe sampling, clipping into short segments, resizing, normalisation
Model consumessingle frames, clips X∈RT×H×W×CX \in \mathbb{R}^{T \times H \times W \times C}, or spatio-temporal patches (tokens)
Enablestraffic analysis, sports analytics, monitoring of industrial processes, content understanding

A video is not stored as a stack of images. Modern codecs store a few complete frames and describe most others as changes relative to their neighbours, which is why video files are so much smaller than their decoded content. A four-second test clip of 64 × 48 pixels at 25 frames per second:

import av
 
n = 0
with av.open("clip.mp4") as container:
    for frame in container.decode(video=0):        # one decoded frame at a time
        arr = frame.to_ndarray(format="rgb24")
        n += 1
print("decoded frames", n, "each", arr.shape)
decoded frames 100 each (48, 64, 3)

The file weighs 3,801 bytes; the 100 decoded frames would occupy 921,600. The ratio is extreme because the test clip is a white square moving on black, but the direction holds for real footage. One minute of 1080p video at 30 frames per second, decoded to 8-bit RGB, is 60×30×1080×1920×3≈11.260 \times 30 \times 1080 \times 1920 \times 3 \approx 11.2 GB.

That arithmetic explains why video pipelines almost never load a whole video into one 4D tensor. They decode as a stream, sample frames (every fourth frame, a few frames per second), cut short clips, and discard what they do not need. A frame-by-frame detector consumes one image at a time. An action-recognition model consumes clips of, say, 16 frames:

with av.open("clip.mp4") as container:
    frames = [f.to_ndarray(format="rgb24")
              for i, f in enumerate(container.decode(video=0)) if i % 4 == 0 and i < 64]
clip = np.stack(frames)
print("clip", clip.shape)
clip (16, 48, 64, 3)

Transformer-based video models go one step further and cut clips into small spatio-temporal blocks, each becoming a token (Arnab et al., 2021). The same video can therefore be a stream of images, a set of clips or a sequence of tokens, depending on what the model needs to notice: an object, a motion, or an event that unfolds over seconds.

Text

QuestionAnswer
In the worldlanguage: words, symbols, code, numbers written as characters
Stored asUTF-8 text files, JSON, HTML, PDF, database fields
Decoded bya character decoder (bytes to Unicode), often after extraction from PDF or HTML
In memorya string: a sequence of Unicode characters
Transformed bynormalisation, splitting into tokens, mapping tokens to integer IDs
Model consumestoken IDs (t1,…,tn)(t_1, \dots, t_n), turned inside the model into vectors X∈Rn×dX \in \mathbb{R}^{n \times d}
Enablessearch, retrieval, classification, question answering, LLM systems

Text is the case where "it becomes a matrix" is most misleading, because there are three different numerical objects and people often call all of them embeddings.

The first step converts characters into tokens: units from a fixed vocabulary learned on a training corpus. Frequent words are single tokens; rare words are split into pieces, which lets a finite vocabulary cover any input (Sennrich et al., 2016). Each token is then replaced by its ID, an integer index into the vocabulary. With the tokenizer of BERT-base (uncased):

from tokenizers import Tokenizer
 
tok = Tokenizer.from_pretrained("bert-base-uncased")
enc = tok.encode("Tokenizers don't read words; they read pieces of words.")
print(enc.tokens)
print(enc.ids)
print("vocabulary size", tok.get_vocab_size())
['[CLS]', 'token', '##izer', '##s', 'don', "'", 't', 'read', 'words', ';', 'they', 'read', 'pieces', 'of', 'words', '.', '[SEP]']
[101, 19204, 17629, 2015, 2123, 1005, 1056, 3191, 2616, 1025, 2027, 3191, 4109, 1997, 2616, 1012, 102]
vocabulary size 30522

Nine words became 17 IDs. "Tokenizers" was split into three pieces (## marks a piece that continues a word), the apostrophe of "don't" became its own token, and two special tokens were added at the edges because BERT was trained with them. The IDs carry no meaning by themselves: 2616 is not "close" to 2617 in any useful sense. They are addresses.

The second object is the embedding. Inside the model, each ID selects one row of a learned table E∈RV×dE \in \mathbb{R}^{V \times d}, where VV is the vocabulary size and dd the width of the model. For BERT-base, d=768d = 768; for BERT-large it is 1024 (Devlin et al., 2018). Other models use other widths, so there is no "standard" embedding size.

import torch
 
ids = torch.tensor([enc.ids])
emb = torch.nn.Embedding(tok.get_vocab_size(), 768)   # untrained table with BERT-base dimensions
print("ids", tuple(ids.shape), ids.dtype, "-> vectors", tuple(emb(ids).shape))
ids (1, 17) torch.int64 -> vectors (1, 17, 768)

This table is randomly initialised, so its vectors mean nothing yet; only the shapes are real. In a trained model, the rows are whatever helped the model reduce its training loss. They often place related tokens near each other, but they are not a dictionary of meanings, and their geometry depends on the model, the data and the objective.

The third object is the contextual representation: the vectors that come out of the model's layers. At the input, the two occurrences of "read" in our sentence receive the same row of EE (BERT adds a position embedding to each, so their input vectors already differ by position). After a few transformer layers, each occurrence has a different vector that depends on the surrounding tokens (Vaswani et al., 2017). When a retrieval system stores "sentence embeddings", it usually stores a vector computed from these contextual representations, not from the input table.

The distinction has practical consequences. Token counts drive cost and context limits, and they differ across tokenizers and languages. Embeddings from two different models live in unrelated spaces and cannot be compared. And a tokenizer is part of the model: swapping it silently breaks everything downstream.

Graphs

QuestionAnswer
In the worldentities and their relations: accounts and transfers, users and items, proteins and interactions
Stored asedge lists (CSV), GraphML, JSON, graph databases, relational tables with foreign keys
Decoded bya parser or driver that builds a graph object (NetworkX, a graph database client)
In memoryadjacency lists, a sparse adjacency matrix, or an edge index plus feature tables
Transformed bynode and edge feature extraction, sampling of neighbourhoods, normalisation
Model consumesnode features X∈R∣V∣×FX \in \mathbb{R}^{\lvert V \rvert \times F} plus connectivity, for message passing
Enablesfraud detection, recommendation, network analysis, reasoning over knowledge graphs

A graph G=(V,E)G = (V, E) is a set of nodes VV and a set of edges EE between them. It has no natural grid and no natural order: renumbering the nodes describes exactly the same graph. That is the deepest difference from images and text.

The textbook representation is the adjacency matrix AA, a ∣V∣×∣V∣\lvert V \rvert \times \lvert V \rvert table where Aij=1A_{ij} = 1 if nodes ii and jj are connected. It is a fine mathematical object and usually a poor data structure. Here is Zachary's karate club, a classic 34-member social network shipped with NetworkX:

import networkx as nx
 
G = nx.karate_club_graph()
print("nodes", G.number_of_nodes(), "edges", G.number_of_edges())
 
edges = np.array(list(G.edges())).T
edge_index = np.concatenate([edges, edges[::-1]], axis=1)   # both directions
print("edge_index", edge_index.shape)
 
A = nx.to_scipy_sparse_array(G, weight=None)
print("sparse adjacency", A.shape, "stored values", A.nnz)
print("dense zeros", f"{(A.toarray() == 0).mean():.1%}")
nodes 34 edges 78
edge_index (2, 156)
sparse adjacency (34, 34) stored values 156
dense zeros 86.5%

Even in this tiny, tightly knit group, 86.5% of the dense matrix is zeros. Large real networks are far sparser: each node connects to a tiny fraction of the others, while a dense matrix grows with the square of the number of nodes. A million nodes would need a trillion cells, almost all of them zero. Practical systems store only the edges. PyTorch Geometric, a widely used library for graph neural networks, represents connectivity as an edge_index of shape [2,∣E∣][2, \lvert E \rvert] (each column is a source and a target, an undirected edge being listed in both directions) and node attributes as a feature matrix of shape [∣V∣,F][\lvert V \rvert, F] (PyG documentation).

Models then compute on that structure by message passing: each node updates its vector by aggregating the vectors of its neighbours, layer after layer (Kipf & Welling, 2017; Gilmer et al., 2017). Other approaches compress each node into a fixed-size embedding learned from random walks (Grover & Leskovec, 2016), which a classical model can then use as ordinary features. The graph never "becomes a matrix" in the sense the slogan suggests. It becomes a few arrays plus a rule for moving information along edges.

Graphs are not the only irregular case. A lidar scan is a set of 3D points with no order at all, and models such as PointNet are designed to give the same answer whatever order the points arrive in (Qi et al., 2017).


The whole picture in one table

Data typeTypical storageDecoding / loadingIn memoryPossible model representationsExample task
Table, time seriesCSV, Parquet, SQL tablesparser, columnar reader, drivertyped columns; float array T×FT \times Ffeature vectors; windows N×T×FN \times T \times F; categorical embeddingsforecasting, anomaly detection
ImageJPEG, PNG, TIFF, DICOMimage codecuint8 array H×W×CH \times W \times C (C = 1, 3 or 4)normalised N×C×H×WN \times C \times H \times W; patch sequenceinspection, classification, segmentation
AudioWAV, FLAC, MP3audio decoder, resamplerwaveform, nn samples per channelwaveform; spectrogram; log-Mel; learned featuresspeech recognition, sound event detection
VideoMP4, MKV, WebMvideo decoder, often hardwarestream of frames H×W×CH \times W \times Cframes; clips T×H×W×CT \times H \times W \times C; spatio-temporal tokensaction recognition, traffic analysis
TextUTF-8, JSON, HTML, PDFcharacter decoding, extractionUnicode stringtoken IDs; embeddings n×dn \times d; contextual vectorssearch, classification, QA
Graphedge lists, GraphML, graph DBparser, driveradjacency lists; sparse matrixedge index + node and edge features; node embeddingsfraud detection, recommendation

Read the table by columns, not by rows. The storage column is about engineering constraints. The memory column is about libraries. The model column has several entries per row, and that is the point: the right one depends on the question asked of the data.


The deeper pattern: representation is a decision

"Everything becomes a matrix" describes the end of the pipeline as if it were fixed. In practice, the representation is chosen, and the choice changes what a system can do.

  • What becomes visible. A pure tone and its first harmonic are two lines in a spectrogram and an unreadable wiggle in a waveform. A daily cycle in sensor data is obvious once "hour of day" is a feature and invisible in a raw timestamp.
  • What is preserved. Resampling to 16 kHz discards everything above 8 kHz. A Mel spectrogram discards phase. Grayscale conversion discards colour, which is fine for reading a serial number and fatal for grading fruit.
  • What it costs. Our two seconds of audio are 88,200 numbers as a native waveform and 128×87=11,136128 \times 87 = 11,136 as a Mel spectrogram. A dense adjacency matrix grows with the square of the number of nodes; an edge list grows with the number of edges.
  • Which architecture fits. Grids suit convolutions. Sequences suit recurrent networks and transformers. Sets need models that ignore order. Graphs need message passing. Pick the representation and you have largely picked the family of models.
  • How easily the model learns. A representation that exposes the right structure saves the model from learning it from scratch. Learning the representation from rawer inputs is possible, but it usually costs data: the Vision Transformer, which starts from plain image patches, needed large-scale pre-training to match convolutional networks that build in assumptions about images (Dosovitskiy et al., 2020).

This section is a synthesis rather than a single sourced fact, but each item above is illustrated by an output earlier in the article. Together they explain why so much of the work in applied machine learning happens before the first training run.


What the hardware wants

Why insist on numbers at all? Because of how processors spend their time.

A CPU has a few powerful cores and vector units that apply one instruction to several numbers at once. NumPy's speed comes from running compiled loops over contiguous arrays and using those vector instructions, which is why the object array in the tabular example is a dead end: it is an array of pointers to Python objects, and every operation has to go back through the interpreter.

GPUs push the same idea much further, with thousands of threads working in parallel. They are general-purpose parallel processors (CUDA exists precisely to program them for arbitrary computations), not machines that "only understand matrices". What makes them so effective for deep learning is that the core layers of neural networks reduce largely to dense matrix multiplication: NVIDIA's own performance guide describes general matrix multiplications as a fundamental building block of fully-connected, recurrent and convolutional layers, and recent GPUs include Tensor Cores dedicated to them (NVIDIA). Specialised accelerators go further still: the first Google TPU was built around a matrix unit of 65,536 multiply-accumulate cells (Jouppi et al., 2017).

So the accurate statement is not "GPUs need matrices". It is this: the representation of data determines which computations can be performed efficiently. A batch of same-sized images maps onto dense, regular arithmetic that hardware handles extremely well. A graph with irregular neighbourhoods or a set of variable-length sentences does not, which is why libraries pad sequences to a common length, group them into batches, and use dedicated sparse operations for graphs. Part of designing a representation is negotiating between what describes the data faithfully and what the hardware can compute quickly.


A mental model to keep

A file is a storage format. A decoder turns it back into values, sometimes making choices you did not notice. A numerical representation in memory makes computation possible. A model consumes a representation designed for its architecture, and two models may want very different ones.

When a new kind of data lands on your desk, ask the seven questions in order, print the shape and the dtype at every step, and check the defaults of every loader. The work is rarely "feeding data to AI". It is deciding which numerical picture of reality will make the computation you care about possible, cheap and accurate.

If you work mostly with tables, the companion article Beyond Syntax covers the operations that act on them once they are in memory.


References

  1. Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lučić, M., & Schmid, C. (2021). ViViT: A Video Vision Transformer. ICCV 2021.
  2. Baevski, A., Zhou, H., Mohamed, A., & Auli, M. (2020). wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations. NeurIPS 2020.
  3. Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL 2019.
  4. Dosovitskiy, A., et al. (2020). An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. ICLR 2021.
  5. Gilmer, J., Schoenholz, S. S., Riley, P. F., Vinyals, O., & Dahl, G. E. (2017). Neural Message Passing for Quantum Chemistry. ICML 2017.
  6. Grover, A., & Leskovec, J. (2016). node2vec: Scalable Feature Learning for Networks. KDD 2016.
  7. Jouppi, N. P., et al. (2017). In-Datacenter Performance Analysis of a Tensor Processing Unit. ISCA 2017.
  8. Kipf, T. N., & Welling, M. (2017). Semi-Supervised Classification with Graph Convolutional Networks. ICLR 2017.
  9. Qi, C. R., Su, H., Mo, K., & Guibas, L. J. (2017). PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation. CVPR 2017.
  10. Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., & Sutskever, I. (2022). Robust Speech Recognition via Large-Scale Weak Supervision. ICML 2023.
  11. Sennrich, R., Haddow, B., & Birch, A. (2016). Neural Machine Translation of Rare Words with Subword Units. ACL 2016.
  12. Vaswani, A., et al. (2017). Attention Is All You Need. NeurIPS 2017.
  13. Zachary, W. W. (1977). An Information Flow Model for Conflict and Fission in Small Groups. Journal of Anthropological Research, 33(4), 452–473. Available in NetworkX as karate_club_graph.
  14. Documentation: NumPy, pandas, Pillow, librosa 0.11, PyAV, Hugging Face Tokenizers, PyTorch, NetworkX, PyTorch Geometric.
All posts