Phase 4 · Advanced C++ & Systems ProgrammingModule 23~64 min read

Concurrency & the C++ Memory Model

Coordinate threads and asynchronous results while preventing races, deadlocks, and invalid memory-order assumptions.

What you'll learn

Concurrency permits progress on multiple tasks, but shared mutation introduces data races, deadlocks, visibility rules, and cancellation. You will coordinate threads through ownership, locks, conditions, futures, and atomics—starting from the simplest correct design.

By the end, you'll be able to:

  • Start and stop threads with automatic joining
  • Protect invariants with mutexes, scoped locks, and condition variables
  • Return asynchronous results through futures and promises
  • Explain races, happens-before, atomic ordering, and thread-safe design

Threads, joining, and cancellation

A thread executes a callable concurrently with its creator. A std::thread must be joined or detached before destruction; detachment makes lifetime and shutdown difficult. C++20 std::jthread joins automatically and supplies cooperative stop tokens when the callable accepts one.

jthread.cpp
#include <chrono>
#include <iostream>
#include <stop_token>
#include <thread>

int main() {
    std::jthread worker{[](std::stop_token stop) {
        int cycle{};
        while (!stop.stop_requested()) {
            std::cout << "cycle " << ++cycle << '\n';
            std::this_thread::sleep_for(std::chrono::milliseconds{10});
        }
    }};

    std::this_thread::sleep_for(std::chrono::milliseconds{25});
    worker.request_stop();
} // jthread joins
Exact cycle count and output timing are intentionally nondeterministic.

Key idea

Cancellation is a request, not forced termination. Long-running work must check the token at safe points and leave every invariant intact before returning.

Mutexes and invariant protection

A mutex establishes exclusive access to shared state. Protect the entire invariant, not just individual variables. RAII lock types release on every exit path. std::scoped_lockcan acquire multiple mutexes with a deadlock-avoidance algorithm.

transfer.cpp
#include <mutex>

struct Account {
    int balance{};
    std::mutex mutex;
};

bool transfer(Account& from, Account& to, int amount) {
    if (&from == &to || amount <= 0) return false;
    std::scoped_lock lock{from.mutex, to.mutex};
    if (from.balance < amount) return false;
    from.balance -= amount;
    to.balance += amount;
    return true;
}
  • Keep critical sections short but large enough to preserve the full invariant
  • Never call unknown user code while holding an internal lock
  • Use one consistent global lock order or scoped_lock for multiple mutexes
  • Document which mutex protects each shared field

Watch out

A data race on ordinary memory is undefined behavior. “It usually writes one machine word” is not synchronization and does not establish visibility.

Condition variables

A condition variable lets a thread sleep until shared state may satisfy a predicate. Waiting releases the mutex atomically and reacquires it before returning. Always use the predicate form because wakeups can be spurious and another thread may consume the condition first.

blocking_queue.cpp
#include <condition_variable>
#include <mutex>
#include <optional>
#include <queue>

class Queue {
public:
    void push(int value) {
        {
            std::lock_guard lock{mutex_};
            values_.push(value);
        }
        ready_.notify_one();
    }

    std::optional<int> pop() {
        std::unique_lock lock{mutex_};
        ready_.wait(lock, [this] { return closed_ || !values_.empty(); });
        if (values_.empty()) return std::nullopt;
        int value{values_.front()};
        values_.pop();
        return value;
    }

private:
    std::mutex mutex_;
    std::condition_variable ready_;
    std::queue<int> values_;
    bool closed_{};
};
A complete queue would add a close() operation that sets closed_ and notifies all waiters.

Futures and asynchronous results

A future represents a result that will become ready. std::promise publishes a value or exception to its future; std::packaged_task connects a callable to that shared state. std::async is convenient, but its launch policy should be explicit when deferred versus concurrent execution changes behavior.

async.cpp
#include <future>
#include <iostream>
#include <numeric>
#include <vector>

long long sum(std::vector<int> values) {
    return std::accumulate(values.begin(), values.end(), 0LL);
}

int main() {
    auto result{std::async(std::launch::async, sum,
                           std::vector<int>{1, 2, 3, 4, 5})};
    std::cout << "working...\n";
    std::cout << result.get() << '\n';
}
The two lines are ordered here because get() occurs after printing; the computation timing is not observable.

Note

future::get() returns or rethrows the stored exception and can be called once. Use shared_future only when several consumers truly need the same result.

Races, atomics, and happens-before

The memory model defines when side effects from one thread become visible to another. Mutex unlock/lock pairs, thread start/join, and suitable atomic operations create happens-before relationships. std::atomic<T> makes operations on one value indivisible, but it does not automatically protect a multi-field invariant.

OrderingUse
relaxedAtomicity only; counters with no publication relationship
acquireLater operations observe a matching release
releaseEarlier operations become visible to a matching acquire
acq_relRead-modify-write with both roles
seq_cstStrong default total order for sequentially consistent atomics
atomic_counter.cpp
#include <atomic>
#include <iostream>
#include <thread>
#include <vector>

int main() {
    std::atomic<int> completed{};
    std::vector<std::jthread> workers;
    for (int i{}; i < 4; ++i) {
        workers.emplace_back([&completed] {
            // perform independent work
            completed.fetch_add(1, std::memory_order_relaxed);
        });
    }
    workers.clear(); // joins all jthreads during destruction
    std::cout << completed.load(std::memory_order_relaxed) << '\n';
}

Watch out

Use the default sequentially consistent ordering until a proven performance need and a written happens-before argument justify weaker ordering. “Seems safe” is not a memory-model proof.

Thread-safe design

The easiest shared state to synchronize is state that is not shared. Partition immutable input, transfer ownership through queues, aggregate results after joining, and expose thread-safe abstractions rather than naked mutexes. Parallel algorithms help only when work is large, independent, supported by the implementation, and safe under unspecified scheduling.

  • Prefer immutable data and message passing
  • Give each mutable object one synchronization policy
  • Design shutdown and cancellation before starting background work
  • Test with ThreadSanitizer where the platform supports it
  • Measure throughput, latency, contention, and oversubscription

Recap & quick check

Key takeaways

  • jthread joins automatically and supports cooperative stop requests.
  • Mutexes protect invariants; RAII lock objects make release exception-safe.
  • Condition-variable waits require a mutex-protected predicate and may wake spuriously.
  • Atomics provide indivisible operations, while ordering defines cross-thread visibility.
  • Ownership transfer and immutable partitioning are simpler than broad shared mutation.

Quick check

1. What does jthread add over thread?

2. Why use the predicate form of condition_variable::wait?

3. Does one atomic field protect a multi-field invariant?

4. What does relaxed ordering provide?

Next: Module 24 — Coroutines & Asynchronous Design, where suspension expresses workflows without blocking a thread.