Quantitative Methods · Reading 11

Machine Learning in Finance

CFA Level I · Quantitative Methods · Reading 11: Introduction to Financial Data Science · about 32 min

What you'll learn

Module 11.1

Introduction to Financial Data Science

This reading describes how financial data science draws on big data: its types, sources and characteristics, and how the data are processed and visualized. It then explains data mining and machine learning (supervised, unsupervised and reinforcement), overfitting and underfitting, the split into training, validation and test datasets, neural networks, and natural language processing and large language models.

LOS 11.a — Big data, machine learning and AI in finance

What financial data science is for

Financial data science applies both quantitative and qualitative analysis to generate insight into a specific financial question, for example how customers behave or where market prices may move next. Among its main uses in finance are (1) evaluating investment opportunities, (2) optimizing portfolios and (3) mitigating risk. Fintech is where finance and technology meet. Of most interest in investment management are the fintech developments that aid quantitative analysis. The word is also used more broadly for technology developments applied to financial services and for the companies that build them.

The objective is the insight itself. Data-quality work, described below under the characteristics of big data and under curation, supports that objective but is not the objective.

Big data: types and sources

Big data is the broad term for all the potentially useful information being generated. By format it is:

TypeWhat it looks likeExamples
Structured dataRows and columns in a table or databaseFinancial statement line items, daily closing prices
Semi-structured dataNot a table, but tagged so it can be stored in an organized wayHTML code of a web page
Unstructured dataNo predefined organization; usually needs ML/AI or special code to useSocial media posts, voice recordings, pictures, sensor output

Data come from traditional sources (markets, company financial reports, regulatory filings, government statistics) or alternative data sources:

  • Individuals: social media activity, clicks, searches, time spent on web pages (mostly unstructured and growing fast).
  • Business processes: transaction and corporate records such as sales data or banking records (usually structured). They can act as leading or real-time indicators, whereas quarterly reports give a lagging view.
  • Sensors: smartphones, connected cars, smart buildings, satellites. The network of these devices is the Internet of Things (IoT); it produces the largest volumes, mostly unstructured.

Buying alternative data or collecting it by web scraping (automated extraction from websites) raises legal and ethical issues: some of what is captured may be personal information protected by privacy rules.

The "Vs" of big data

Key concept

CharacteristicQuestion it answersVocabulary to recognize
VolumeHow much data?megabytes → gigabytes → terabytes → petabytes
VelocityHow fast are data communicated?low latency = real time; high latency = periodic or delayed
VarietyIn what structures do the data exist?structured, semi-structured, unstructured
VeracityAre the data reliable and credible?fake posts, errors, doubtful sources

A terabyte is 1,000 gigabytes and a petabyte is 1,000 terabytes. Faster data generate more of it: volumes rise from batch collection (megabytes) through periodic (gigabytes) and near-real-time (terabytes) to real-time feeds (petabytes). More complex structures also need more storage.

The first three are the classic 3 Vs (volume, velocity, variety); veracity is the fourth V that matters especially for financial data. Other data-quality concerns are outliers, bad or missing data, sampling biases, and whether the amount collected is sufficient and appropriate for the intended use.

Data science: processing and visualization

Data science is the field concerned with extracting information from big data. Its processing methods:

MethodWhat happensDo not confuse with
CaptureCollecting data and converting them into a usable formTransfer
CurationEnsuring data quality: fixing errors, adjusting for bad or missing dataCapture
StorageArchiving data and accessing them later; choice depends on structure and latencySearch
SearchQuerying stored data to find what is neededStorage
TransferMoving data from the source or storage medium to the analytical toolCapture

A data lake is a central repository that keeps structured, semi-structured and unstructured data in their raw, original format. Distributed computing frameworks let data be processed and analyzed on many machines at once.

Visualization: charts and graphs suit structured data; less structured data call for word clouds (show how often words occur in a text sample), heat maps, 3D graphics and mind maps (show logical relations between concepts).

Typical inputs include share prices and trading volumes, asset returns, interest rates, exchange rates, accounting figures and macroeconomic series. The work calls for expertise in finance as well as in data handling, because much of the data is sequential, time dependent and sometimes cyclical. Financial data bring special problems: noise, nonstationarity, interdependencies, nonlinearity, fat tails (extreme events more likely), very high-frequency data in trading, and too little data in some markets (e.g., emerging markets). Regulated firms must meet strict rules on data management, risk assessment and model validation; firms outside financial regulation still have to comply with business law and their contracts, although they have more freedom in how they handle data. Data privacy, thorough model testing and human review of model decisions matter for every firm.

Data mining, AI and machine learning

Data mining searches large datasets (usually structured) for patterns that can be used to predict outputs; pattern identification rests on conditional probabilities (e.g., Bayes' theorem). Its key risk is overfitting.

Artificial intelligence (AI) refers to computer systems that simulate human thinking. Machine learning (ML) is a major subcategory of AI: an algorithm receives input data, without any assumption about how they are distributed, possibly with target outputs, and learns without human help to model the outputs or to find patterns.

Key concept

ApproachData givenWhat the machine doesTypical finance use
Supervised learningLabeled inputs and outputsLearns to model outputs from inputs, then predicts on new dataForecasting company or stock performance, default prediction
Unsupervised learningUnlabeled inputs, no outputsDescribes the structure of the data and finds patterns; reduces dimensionalityCustomer segmentation, anomaly detection
Reinforcement learningFeedback on its own actionsLearns by trial and error from rewards and penaltiesPortfolio optimization (sequential decisions)

A supervised model trained on a large, reliable dataset should predict well but may not adapt to new situations. Reinforcement learning is costly: designing a reward system that actually improves the model can take much training and many resources. Whatever the method, people with subject expertise are needed to choose the model, clean the inputs, limit bias and interpret the output.

Deep learning uses many layers of neural networks to find patterns of increasing complexity and may use supervised, unsupervised or semi-supervised learning (image and speech recognition, natural language processing).

Overfitting vs. underfitting

Key concept

OverfittingUnderfitting
Model isToo complexNot complex enough
ErrorTreats noise as a true patternTreats true parameters as noise
ResultSpurious relationships, poor generalization to new dataMisses actual patterns and relationships

Other ML issues: results can be a "black box" (hard to explain), and data leakage occurs when information from outside the training dataset influences what the model learns.

Training, validation and test datasets

  • Training dataset: the model looks for relationships here; it is the largest set, about 60%–80% of the data (upper end for large datasets).
  • Validation dataset: used to refine (tune) the relationships found in training.
  • Test dataset: used last, to judge the model's predictive ability.
    The validation and test sets share the remaining 20%–40% equally. Observations are normally assigned at random, which reduces bias, overfitting and data leakage. Time series data must keep their chronological order instead:
  • Time-based split: the oldest 60%–80% of observations go to training, the next 10%–20% to validation and the most recent 10%–20% to testing.
  • Rolling-window validation: train on one window, validate on a later period, then shift the window forward and repeat.
    Either way, the span covered by validation plus test data should match the intended forecast horizon.

Example. A credit model has 48,000 randomly ordered loan records and uses a 75% / 12.5% / 12.5% split. Training gets 36,000 records, validation 6,000 and test 6,000. If instead the data were 12 years of monthly returns (144 months) and a time-based split with the same proportions were used, months 1–108 would train the model, months 109–126 would validate it and months 127–144 would test it.

Neural networks, NLP and LLMs

Neural networks suit nonlinear relationships with many interdependencies: an input layer, several hidden layers of weighted nodes (neurons), and an output layer. They can learn in a supervised or unsupervised way, and each neuron's weight reflects its importance. They are often trained by backpropagation (a form of supervised learning): a forward pass produces a prediction that is compared with the actual outcome, the loss is computed, the loss gradient shows how weight changes affect the error, and gradient descent adjusts the weights step by step to minimize the loss.

  • Generative adversarial networks (GANs) pit a generator network against a discriminator network to create realistic synthetic data (e.g., plausible market scenarios for stress tests, trained on historical market data). Because GANs can also produce convincing fake material, their output needs careful review.
  • Variational autoencoders (VAEs) compress data and then reconstruct them using encoder and decoder networks. They are probabilistic generative models, used to reduce dimensionality and spot anomalies.
  • Copulas (e.g., the Gaussian copula) are a more traditional mathematical tool for nonlinear relationships that uses marginal probability distributions to capture dependencies between random variables. Exam convention: the Cholesky decomposition is given as an example of a copula. Current practice: the Cholesky decomposition is a matrix factorization, used to generate the correlated normal draws in a Gaussian-copula simulation.

Natural language processing (NLP) lets computers interpret human language: speech recognition, translation, scanning employee communications for compliance, and detecting sentiment shifts in research reports or earnings-call transcripts. Text analytics works mainly with unstructured material, such as filings and research reports or recordings of earnings calls and surveys. NLP converts that material into usable data much more quickly than a person could and searches it for patterns; lexical analysis, for example, measures how often each word appears. Large language models (LLMs) are generative AI that produce human-like text, using neural networks to read context and meaning in qualitative inputs (nonfinancial inputs such as election polls can matter too). Their output is text such as a summary or a forecast based on tone and sentiment. They are typically trained on a very large dataset through self-supervised learning, an unsupervised approach in which the model generates the labels itself from raw data. They are not built to catch errors in financial data, so outputs need expert review and should be used alongside traditional data-driven models.

Common tools are Python (fintech apps), R (statistics), Java (runs across platforms) and C/C++ (high-speed algorithmic trading). Common databases are SQL (structured data, on a server), SQLite (structured data, not on a server; common in mobile apps) and NoSQL (unstructured data).

Common exam traps

  • Business-process records such as sales data are usually structured, yet they count as alternative data. Structure does not decide whether a source is traditional.

Exam shortcuts

  • Classify a data source by where it comes from, not by its structure: records from individuals, business processes and sensors are alternative data even when they are structured.
  • Match the type of machine learning to what the algorithm is given: labeled inputs and outputs mean supervised learning, unlabeled inputs alone mean unsupervised learning, and rewards and penalties for its own actions mean reinforcement learning.

Bottom line

  • Financial data science applies quantitative and qualitative analysis to gain insight into a specific financial question, with main uses in evaluating investment opportunities, optimizing portfolios and mitigating risk.
  • Big data can be structured (rows and columns), semi-structured (tagged, such as HTML code) or unstructured (social media posts, voice recordings, pictures, sensor output), and comes from traditional sources or from alternative sources: individuals, business processes and sensors.
  • The characteristics of big data are volume, velocity (low latency means real time), variety of structures and, for financial data especially, veracity, meaning whether the data are reliable and credible.
  • Data science processes data through capture, curation (ensuring data quality), storage, search and transfer, and a data lake keeps structured, semi-structured and unstructured data in their raw format.
  • Data mining searches large datasets for patterns to predict outputs, while machine learning, a subcategory of AI, learns from input data without assumptions about their distribution and without human help.
  • Supervised learning uses labeled inputs and outputs, unsupervised learning finds structure in unlabeled inputs, and reinforcement learning learns by trial and error from rewards and penalties.
  • Overfitting comes from a model that is too complex and treats noise as a true pattern; underfitting comes from one that is not complex enough and treats true parameters as noise.
  • The training dataset is the largest, about 60%–80% of the data, with validation and test sets splitting the rest equally; observations are normally assigned at random, but time series data keep their chronological order through a time-based split or rolling-window validation.

Quick check

Question 1Core

A new analyst at Ottery Capital drafts three glossary entries for the firm's research team. Which entry about fintech is most accurate?

Show answer and explanation

Correct answer: C

Fintech refers to the meeting of the finance and technology worlds. In investment management the developments of most interest are those that aid quantitative analysis, such as big data tools and machine learning. The term is also used more broadly for technology applied to financial services and for the firms that build it.

Why the other options are wrong

  • A. Consumer payment and banking apps are one use of fintech. Limiting the term to them leaves out the investment tools, such as machine learning and big data processing, that matter most to analysts.
  • B. Financial data science is the quantitative and qualitative analysis of financial information to gain insight into a financial question. Fintech is the technology that supports such work and other financial services, so the two terms are related but different.

Key takeaway Fintech is where finance meets technology; financial data science is the analysis that some of that technology makes possible.

Practice Questions

Question 2Core

Voice recordings of earnings calls, satellite images of shipping ports and posts on social media platforms are all examples of:

Show answer and explanation

Correct answer: A

Unstructured data lack a predefined, tabular organization; examples include social media content, voice recordings, pictures and sensor output. Because they are hard for a person to handle efficiently, analyzing them usually requires ML algorithms, AI or specialized code.

Why the other options are wrong

  • B. Structured data are held in tables of rows and columns in a database (e.g., balance sheet data, market price records). Audio, images and social media posts are not in that form.
  • C. Semi-structured data cannot be stored in tables but are tagged so they can be stored in an organized way, such as HTML code. The examples given have no such tagging structure.

Key takeaway Social media, voice, pictures, video and sensor data are unstructured.

Question 3Core

A data team at Larchmont Analytics is documenting the steps it uses to process data. Which of the following statements about these data processing methods is most accurate?

Show answer and explanation

Correct answer: B

Curation is the processing step that assures data quality: errors are corrected and bad or missing observations are adjusted for before analysis. The other two statements attach the description of one processing step to the name of a different step.

Why the other options are wrong

  • A. Archiving data and accessing them later describes storage. Search means querying data that are already stored to locate the information needed.
  • C. Moving data from their source or a storage medium to the analytical tool is transfer. Capture is the collection of data and their conversion into a usable form in preparation for analysis.

Key takeaway Learn the five processing methods as pairs that are easy to mix up: capture vs. transfer, and storage vs. search. Curation is always about data quality.

Question 4Core

A lender at Brackenridge Credit trains a model on records of past borrowers, each tagged with the borrower's financial ratios and whether the borrower eventually defaulted. The model learns to map the ratios to the default outcome and is then applied to new loan applicants. This machine learning process is best described as:

Show answer and explanation

Correct answer: B

Supervised learning works with labeled data: every input record comes with its known output, the machine learns how the outputs relate to the inputs, and the fitted model is then used on new data to forecast outcomes. Here the ratios are the inputs, the default outcome is the labeled output, and the model is applied to new applicants, so it is supervised learning.

Why the other options are wrong

  • A. Unsupervised learning receives unlabeled inputs only, with no target outputs; the machine uncovers structure and patterns on its own (e.g., customer segmentation, anomaly detection). The lender does supply outputs (default or not).
  • C. Reinforcement learning models learn by trial and error, improving through feedback in the form of rewards or penalties. No reward system is described here.

Key takeaway Labeled inputs and outputs: supervised. No outputs: unsupervised. Rewards and penalties: reinforcement.

This reading has 25 questions in the full bank. Practice all of them.

Key Takeaways