0% found this document useful (0 votes)
17 views125 pages

C++ and Python Integration Techniques

The document provides an overview of a quantitative research developer's work at Tower Research Capital, focusing on the development of low latency trading systems using C++ and Python. It discusses the advantages of using both languages, including performance and ease of use, and details various methods for integrating C++ with Python through APIs and libraries like pybind11 and Boost.Python. Additionally, it covers performance metrics for different implementations and offers insights into software design and binding techniques for efficient data processing.

Uploaded by

quanot
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
17 views125 pages

C++ and Python Integration Techniques

The document provides an overview of a quantitative research developer's work at Tower Research Capital, focusing on the development of low latency trading systems using C++ and Python. It discusses the advantages of using both languages, including performance and ease of use, and details various methods for integrating C++ with Python through APIs and libraries like pybind11 and Boost.Python. Additionally, it covers performance metrics for different implementations and offers insights into software design and binding techniques for efficient data processing.

Uploaded by

quanot
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

A quick bio - Yours truly

● Quantitative research developer at Tower Research Capital


○ High frequency trading firm based out of NYC

● Develop low latency trading systems (C++)


○ Nanoseconds matter

● Develop high throughput research systems (C++ and Python)


○ Data volume in terabytes

● Program analysis research and functional programming in a


past life
● Love performance, software abstractions, and clean APIs
Why Python? Why C++?
Why Python?

● Writing extensive APIs in Python - low boilerplate


● Familiar for domain experts
● Easy to use
○ Amazing interactive support out of the box (IPython)
○ Jupyter notebooks provide a great research environment
○ Very mature open source libraries for various domains

Why C++?

● We’re at CppNow :)
Why C++ and Python
● Why?
○ Avoid reimplementing complex code for Python
○ Performance
○ Back and forth with user’s python code
○ Interoperability with data structures in Python - shared memory space

● How?
○ Python/C API
○ Cython
○ Numba / PyPy
○ Boost::Python
○ Pybind11
○ cppyy
○ Mix and match!
Hello, world
(not really)
Hello, world
Different implementations of a simple increment function

(takes an integer as argument, increments it and returns)

● C++
● Python
● C++ with pybind11
● C++ with boost::python
C++
int increment(int); // defined in a separate translation unit to prevent inlining

int main(int argc, char**argv) {


int max_counter = std::stoi(std::string(argv[1]));
int i = 0;
auto start = high_resolution_clock::now();
while (i < max_counter) {
i = increment(i);
}
auto end = high_resolution_clock::now();
std::cout << "Time taken per call: " << ... << std::endl;
}
Python
def increment(i):
return i + 1

if __name__ == "__main__":
i = 0
max_i = 2**20
if len([Link]) > 1:
max_i = int([Link][1])
start = [Link]()
while i < max_i:
i = increment(i)
end = [Link]()
print("Time taken: {:.2f}ms".format((end - start) * 1000))
print("Time taken per iteration: {:.2f}ns".format((end - start) * 1000000000 / i))
pybind11
int cpp_increment(int i) { return i + 1; }

PYBIND11_MODULE(hello_world_pybind11, m) {
[Link]() = "pybind11 incrementer";

[Link]("increment", &cpp_increment, "Incrementing function");


}

// After compiling above into hello_world_pybind11.so, in python


from hello_world_pybind11 import increment
while i < max_i:
i = increment(i)
Boost::python
int cpp_increment(int i) { return i + 1; }

BOOST_PYTHON_MODULE(hello_world_bpy) {
using namespace boost::python;
def("increment", &cpp_increment, "Incrementing function");
}

// After compiling above into hello_world_bpy.so, in python


from hello_world_bpy import increment
while i < max_i:
i = increment(i)
Wait, what?
Enter the Cpython interpreter
● Reference implementation of the python interpreter
● A C program that interprets python instructions
○ Internally maintains state of your program (objects etc) as C objects
derived from a struct called PyObject.
○ Each new python statement tells it how to modify these PyObjects.
○ State of the program is maintained in these introspectable objects.
Enter the Cpython interpreter
● Loading a .so into this program’s memory space allows us
to interact with its internals.
● We can call methods (from <Python.h>) to change the state of
the program
○ API documented nicely [Link]

● We can add “hooks” into the program to invoke our methods


in certain situations.
● Such a module is called an “extension module”.
Python C API
#include <Python.h>

int cpp_increment(int i) { return i + 1; }

static PyObject* python_increment(PyObject* self, PyObject* args) {


int i;
if(!PyArg_ParseTuple(args, "i", &i)) {
return NULL;
}
return PyLong_FromLong(cpp_increment(i));
}
static PyMethodDef HWMethods[] = {
{"increment", python_increment, METH_VARARGS, "Incrementing function"},
{NULL, NULL, 0, NULL}
};
static struct PyModuleDef hello_world_python = {
PyModuleDef_HEAD_INIT, "hello_world_python", "Simple python module", -1, HWMethods
};
PyMODINIT_FUNC PyInit_hello_world_python(void) {
return PyModule_Create(&hello_world_python);
}
Digression: Jython / IronPython / PyPy
[Link]

Easier to extend

Can import libraries / modules from the host language


Various libraries to “speed up” python

Focus of this talk


Performance
Some perf numbers: per increment runtime
Take these numbers with a grain of salt
Some perf numbers: per increment runtime

Implementati 1 run 10 runs 100 runs 1000 runs Size


on

Python 953ns 143ns 71ns 67ns NA

C++ 85ns 11ns 3ns 1ns 91K*

Python C API 1192ns 214ns 83ns 74ns 14K

pybind11 9536ns 1525ns 634ns 552ns 279K

boost::python 1907ns 309ns 123ns 109ns 237K

*: C++ program contains the std::chrono library to time, others use python to do that
Everything compiled with clang++12 -fPIC -O3
Some perf numbers: per increment runtime
[Link]
Out in the wild case of noticing high overheads in
pysimdjson (they switched from pybind11 to cython)
Some GitHub issues:
[Link]
[Link]
Primary causes seem to be an std::vector to store arguments,
and some extra argument type casting work.

This latency of calling a function may or may not be relevant to everyone. But is good to know.
Classes 101
Classes?
class RowReader : public BinaryListener {
std::vector<Row> rows_;
BinaryReader reader_;
public:
RowReader(const std::string& filename); // somehow populates rows_ vector
void onData(int id, BData data) override; // BinaryListener virtual
std::string getIdName(int id) const;
auto getRows() { // can have variants that filter the rows
return rows_;
}
};
Classes!
BOOST_PYTHON_MODULE(binary_reader_bpy) {
class_<binary_reader::Row>("Row", no_init)
.def("nanotime", &Row::nanotime)
.def("items", &Row::items, return_value_policy<reference_internal_object>());
class_<std::vector<Row>>("cpp::vector<Row>")
.def(vector_indexing_suite<std::vector<Row>>());
class_<RowReader>("RowReader", init<std::string>(arg("filename")))
.def("getRows", &RowReader::getRows)
.def("getIdName", &RowReader::getIdName);
}
Classes!
Classes!

import binary_reader_bpy as mod


reader = [Link]("./testfile")
rows = [Link]()
print("Number of rows: " + str(len(rows)))
for i, row in enumerate(rows):
print("Row " + str(i) + " is " + str(row))
for item in [Link]():
print(" - {}:{}".format([Link]([Link]), [Link]()))
Working example
$ python3 python/[Link]
Number of rows: 2
Row 0 is 1682563241000000000: 4 items
- trade_size:3
- price:4086.2
- trade_size:3
- price:4086.2
Row 1 is 1682563242000000000: 7 items
- trade_size:5
- price:4086.2
- trade_size:5
- price:4086.7
- signal:1.2
- order_size:1
- price:4086.7
Software design with
pybindings
Basic thought process
● Define a user boundary for your library’s API
● Expose methods at that boundary using python binding
● Try putting “hot path” code inside your library
○ Need to design a DSL for the user to be able to leverage your hot
path capabilities written in C++
○ Numpy / Pandas / PyTorch are nice DSLs that provide “hot path” C++
functions

● Can allow python objects for constructors / non-critical


operations
○ Simply for ease of integration
What Functions to expose in the API
● Init some data structures / load some data
● Configure your class’ parameters
● Do a lot of complex work inside a “hot loop”
● Fetch internal state of your data structure
Arguments to functions
● Can be int / float / std::vector<std::string>, ...
○ Most libraries support all these common types’ conversion from
PyObject* and back
○ Should be optimal to let the library do the cast for you

● Can be instances types defined by your module


● Can be PyObject* / boost::python::object
○ boost::python::object is a smart pointer around a PyObject*
○ Only useful to use bpy::object if you’re going to store it
○ Why would we want to do this?
Arguments to functions - PyObjects

MyReader::MyReader(bpy::object filenames) {
std::vector<std::string> filenames_std;
bpy::extract<bpy::list> get_list(filenames);
if (get_list.check()) {
filenames_std = to_std_vector<std::string>(get_list());
} else {
bpy::extract<std::string> get_string(filenames);
if (get_string.check()) {
filenames_std.push_back(get_string());
} else {
// error
}
}
Return values of functions
● Easy enough to return standard types (int, float, ...)
○ Let bindings library do the heavy lifting

● Can return python objects


○ boost::python::dict / boost::python::list are common ones I use
○ boost::python just makes it easy to create refcounted PyObjects
○ Even more optimal if your code just populates the python data
structure instead of doing a two pass

● What about larger objects?


○ Large vectors, matrices?
○ Python libraries aren’t written to operate on std::vector<float>
Return values of functions - numpy

#include <boost/python/[Link]>
namespace bnp = boost::python::numpy;

template <typename T>


inline bnp::ndarray to_boost_vector(const std::vector<T>& v) {
Py_intptr_t shape[1] = {static_cast<long>([Link]())};
bnp::ndarray result = bnp::zeros(1, shape, bnp::dtype::get_builtin<T>());
std::copy([Link](), [Link](), reinterpret_cast<T*>(result.get_data()));
return result;
}
Lifetime of returned values
● Say you return a struct (struct1) containing references to
another struct (struct2)
● Python will manage struct1’s instance as a refcounted
object
○ Note: Python is a reference counted garbage collected language
● Python has no way to know struct1 refers to struct2
internally, and will happily garbage collect struct2
● Classic use for shared pointers, just good to remember
Lifetime of returned values

class Item {
const Reader* reader_;
int id;
double value;
public:
std::pair<std::string, double> getNameValue() {
return {reader_->getName(id), value};
}
};
Lifetime of returned values

class Item {
std::shared_ptr<const Reader> reader_;
int id;
double value;
public:
std::pair<std::string, double> getNameValue() {
return {reader_->getName(id), value};
}
};
A note about “software glue”
● Composing softwares requires a scripting “glue” at times
● The unix philosophy is to spin up lots of processes, and
using pipes to compose softwares
○ Hard to do easy back and forth b/w softwares using OOP
○ Hard to shared typed data
○ Hard to optimize away the last few copies
● Python’s ease of extension => powerful glue
Example 1
Custom binary
data format
Naive approach: Custom binary logging data
● Have existing C++ library to parse data
● Write C++ to serialize data as plaintext on stdout
● Write python library to process data items from stdout
● Write analysis code using python by iterating over all
this data

Too slow! 😞
Enter python bindings
● Re-use C++ library
● Make class that stores / preprocesses data in memory
● Expose API methods to process that data and expose
results in a memory layout that python understands
○ Write DSL to return numpy arrays for analysis
We’ve seen this before
import binary_reader_bpy as mod

print("Module version is " + str(mod.API_VERSION))


reader = [Link]("./testfile")
rows = [Link]()
print("Number of rows: " + str(len(rows)))
for i, row in enumerate(rows):
print("Row " + str(i) + " is " + str(row))
for item in [Link]():
print(" - {}:{}".format([Link]([Link]), [Link]()))
Reader code
class RowReader : public BinaryListener {
std::vector<RowRef> rows_;
public:
RowReader(const std::string& filename); // does the file IO to read file into memory
void onData(int id, BData data, const IdMetadata& metadata) override {
if ([Link] == "timestamp")
rows_.push_back(RowRef{.nanotime_ = std::get<int64_t>(data)});
else
rows_.back()->items_.push_back({id, data});
}
std::string getIdName(int id) const;
std::vector<RowRef> getRows() const;
std::vector<RowRef> getRowsWithSignalGreaterThan(double signal) const;
};
binding code
BOOST_PYTHON_MODULE(binary_reader_bpy) {
class_<binary_reader::Item>("Item", no_init)
.def_readonly("id", &Item::id)
.def("value", &Item::getPyValue);
class_<binary_reader::RowRef>("Row", no_init)
.def("nanotime", &RowRef::nanotime)
.def("items", &RowRef::items, return_value_policy<return_internal_reference>());
class_<std::vector<RowRef>>("cpp::vector<RowRef>").def(vector_indexing_suite<std::vector<RowRef>>());
class_<std::vector<Item>>("cpp::vector<Item>").def(vector_indexing_suite<std::vector<Item>>());
class_<RowReader>("RowReader", init<std::string>(arg("filename")))
.def("getRows", &RowReader::getRows)
.def("getIdName", &RowReader::getIdName);
scope().attr("API_VERSION") = 1;
register_exception_translator<std::runtime_error>(translate_error);
}
(1) Binding code - Basics
BOOST_PYTHON_MODULE(binary_reader_bpy) {
class_<binary_reader::Item>("Item", no_init)
.def_readonly("id", &Item::id)
.def("value", &Item::getPyValue);
class_<binary_reader::RowRef>("Row", no_init)
.def("nanotime", &RowRef::nanotime)
.def("items", &RowRef::items, return_value_policy<return_internal_reference>());
class_<std::vector<RowRef>>("cpp::vector<RowRef>").def(vector_indexing_suite<std::vector<RowRef>>());
class_<std::vector<Item>>("cpp::vector<Item>").def(vector_indexing_suite<std::vector<Item>>());
class_<RowReader>("RowReader", init<std::string>(arg("filename")))
.def("getRows", &RowReader::getRows)
.def("getIdName", &RowReader::getIdName);
scope().attr("API_VERSION") = 1;
register_exception_translator<std::runtime_error>(translate_error);
}
(1) Binding code - Basics
using BData = std::variant<int64_t, double>;

struct Item {
int id;
BData value;
PyObject* getPyValue() const;
bool operator==(const Item& o) const { return [Link] == id && [Link] == value; }
};

// Declaration
class_<binary_reader::Item>("Item", no_init)
.def_readonly("id", &Item::id)
.def("value", &Item::getPyValue);
(2) Binding code -return_value_policy
BOOST_PYTHON_MODULE(binary_reader_bpy) {
class_<binary_reader::Item>("Item", no_init)
.def_readonly("id", &Item::id)
.def("value", &Item::getPyValue);
class_<binary_reader::RowRef>("Row", no_init)
.def("nanotime", &RowRef::nanotime)
.def("items", &RowRef::items, return_value_policy<return_internal_reference>());
class_<std::vector<RowRef>>("cpp::vector<RowRef>").def(vector_indexing_suite<std::vector<RowRef>>());
class_<std::vector<Item>>("cpp::vector<Item>").def(vector_indexing_suite<std::vector<Item>>());
class_<RowReader>("RowReader", init<std::string>(arg("filename")))
.def("getRows", &RowReader::getRows)
.def("getIdName", &RowReader::getIdName);
scope().attr("API_VERSION") = 1;
register_exception_translator<std::runtime_error>(translate_error);
}
(2) Binding code -return_value_policy
struct Row {
std::vector<Item> items_;
const auto& items() const { return items_; }
};

// Declaration:
.def("items", &RowRef::items, return_value_policy<return_internal_reference>());

Let’s try to engineer the best possible semantics for


managing references in python.
(2) Binding code -return_value_policy
.def("items", &RowRef::items, return_value_policy<return_existing_object>()); // UNSAFE!

// Unsafe python code:


def get_first_row_items(reader):
for row in [Link]():
return [Link]()

● Return a PyObject that contains an unmanaged pointer


○ Unsafe, lifetime of “items” PyObject can exceed lifetime of Row
(2) Binding code -return_value_policy
class ItemsVectorWrapper {
const std::vector<Item> items_;
PyObject* row_;
public:
Item operator[](size_t idx) { return items_[idx]; }
ItemsVectorWrapper(const std::vector<Item>& items, PyObject* row) : items_(items), row_(row) {
Py_IncRef(row);
}
~ItemsVectorWrapper() { Py_DecRef(row_); }
};

● Change return value to be a wrapped struct


(2) Binding code -return_value_policy
● Rather verbose and annoying to write

● Luckily, return_value_policy<return_internal_reference> is smart

● It does exactly what we discussed, but on its own


(3) Binding code - vector_indexing_suite
BOOST_PYTHON_MODULE(binary_reader_bpy) {
class_<binary_reader::Item>("Item", no_init)
.def_readonly("id", &Item::id)
.def("value", &Item::getPyValue);
class_<binary_reader::RowRef>("Row", no_init)
.def("nanotime", &RowRef::nanotime)
.def("items", &RowRef::items, return_value_policy<return_internal_reference>());
class_<std::vector<RowRef>>("cpp::vector<RowRef>").def(vector_indexing_suite<std::vector<RowRef>>());
class_<std::vector<Item>>("cpp::vector<Item>").def(vector_indexing_suite<std::vector<Item>>());
class_<RowReader>("RowReader", init<std::string>(arg("filename")))
.def("getRows", &RowReader::getRows)
.def("getIdName", &RowReader::getIdName);
scope().attr("API_VERSION") = 1;
register_exception_translator<std::runtime_error>(translate_error);
}
(3) Binding code - vector_indexing_suite
class_<std::vector<Item>>("cpp::vector<Item>").def(vector_indexing_suite<std::vector<Item>>());

● C++ vectors have value semantics for operator[]


○ myvec[1].mutate() will not change value of myvec[1], unless?

● Python vectors have reference semantics for operator[]


● vector_indexing_suite tries to make c++ vectors behave
like python vectors
(3) Binding code - vector_indexing_suite
class_<std::vector<std::shared_ptr<Row>>>("cpp::vector<std::shared_ptr<Row>>")
.def(vector_indexing_suite<std::vector<std::shared_ptr<Row>>, true>());
class_<RowReader>("RowReader", init<std::string>(arg("filename")))
.def("getRows", &RowReader::getRows, return_value_policy<reference_existing_object>())

● Will memory usage increase after calling getRows?


● No! (I checked 🙂)
○ Does not copy the vector, just exposes an API around it
(4) Binding code - Constructor
BOOST_PYTHON_MODULE(binary_reader_bpy) {
class_<binary_reader::Item>("Item", no_init)
.def_readonly("id", &Item::id)
.def("value", &Item::getPyValue);
class_<binary_reader::RowRef>("Row", no_init)
.def("nanotime", &RowRef::nanotime)
.def("items", &RowRef::items, return_value_policy<return_internal_reference>());
class_<std::vector<RowRef>>("cpp::vector<RowRef>").def(vector_indexing_suite<std::vector<RowRef>>());
class_<std::vector<Item>>("cpp::vector<Item>").def(vector_indexing_suite<std::vector<Item>>());
class_<RowReader>("RowReader", init<std::string>(arg("filename")))
.def("getRows", &RowReader::getRows)
.def("getIdName", &RowReader::getIdName);
scope().attr("API_VERSION") = 1;
register_exception_translator<std::runtime_error>(translate_error);
}
(4) Binding code - Constructor
class_<RowReader>("RowReader", init<std::string>(arg("filename")))
.def("getRows", &RowReader::getRows)
.def("getIdName", &RowReader::getIdName);
register_exception_translator<std::runtime_error>(translate_error);

Some nuances to consider when defining python construction:

● May throw an exception


● May do IO that we may want to parallelize across threads
(4) Binding code - Exceptions
void translate_error(const std::runtime_error& e) {
PyErr_SetString(PyExc_Exception, [Link]());
}
register_exception_translator<std::runtime_error>(translate_error);

● Exceptions are thread local error flags.


● Existing libraries may have ::abort / ::exit calls in
macros
● Use #ifdef macros to switch your exit macros to throw
exceptions instead
(4) Binding code - Exceptions
#ifdef THROW_EXCEPTION_INSTEAD_OF_DYING
#define DIE(msg) MyNs::DieThrower(__FILE__, __LINE__, msg)
#else
#define DIE(msg) MyNs::DieExiter(__FILE__, __LINE__, msg)
#endif
(4) Binding code - Threading
RowReader(const std::string& filename) : reader_(filename) {
reader_.addListener(this);
reader_.readAll();
}

// Python API
with ThreadPoolExecutor(max_workers=10) as executor:
futures = []
for f in [Link][1:]:
[Link]([Link]([Link], f))
readers = [[Link]() for fut in futures]
(4) Binding code - Threading
● Python has a Global Interpreter Lock (GIL)
● A blocking IO call that does not deal with any python
objects should be allowed to release the lock
● Python’s internal IO libraries already do this
● We’ll discuss nuances of this in a later example
(4) Binding code - Threading
class without_gil {
public: /* Source: [Link] */
without_gil() { state_ = PyEval_SaveThread(); }
~without_gil() { PyEval_RestoreThread(state_); }
without_gil(const without_gil&) = delete;
without_gil& operator=(const without_gil&) = delete;
private:
PyThreadState* state_;
};
RowReader::RowReader(const std::string& filename) : reader_(filename) {
without_gil no_gil;
....
}
(4) Binding code - Threading
Reading 3 161 MB files with and without
releasing GIL

# without releasing GIL


$ /usr/bin/time python3 python/[Link] testfile1 testfile2 testfile3
Module version is 1
Finished
10.63user 0.42system 0:11.11elapsed 99%CPU (0avgtext+0avgdata 1509900maxresident)k

# after releasing GIL


$ /usr/bin/time python3 python/[Link] testfile1 testfile2 testfile3
Module version is 1
Finished
10.36user 0.52system 0:04.13elapsed 263%CPU (0avgtext+0avgdata 1509932maxresident)k
Example 2
C++ library with a
dispatcher
Dispatchers?
while (true) {
if (source1.has_data()) {
[Link]();
}
if (source2.has_data()) {
[Link]();
}
}
Dispatchers and callbacks
def print_packet(packet):
print(packet)
my_network_library = MyNetworkLibrary()
my_network_library.register_callback_on_packet(print_packet)
my_network_library.listenAndBlock()

● Callbacks are a common theme in many dispatcher based


applications
● Dispatching may happen in:
○ The main python thread, but blocked
○ A separate thread, thus letting the python thread run
Callbacks syntax
std::vector<boost::python::object> callbacks_;

void registerCallbackOnPacket(boost::python::object callback) {


if (callback != boost::python::object() && !PyCallable_Check([Link]())) {
throw std::runtime_error("Callback has to be callable");
}
callbacks_.push_back(callback);
}

void sendInfoToPython(Info info) {


for (auto& callback : callbacks_) {
callback(info); // TODO: Need to grab GIL before we do this
}
}
Let’s make this complicated, one step at a time
Complexity level 1 : Everything in the same thread
Complexity level 1 : Everything in the same thread
● Very simple, single threaded operation
● No need to think about locking / threads
● Python threads cannot run in the background though
○ Maybe some GUI being run by python
○ Maybe some widgets / resource monitoring code

● Let’s try to add more worker threads to this


Complexity level 2 : Multiple worker threads
Complexity level 2 : Multiple worker threads
● Worker threads cannot call the callback handlers without
acquiring the GIL
● The main watcher thread needs to release the GIL first!
● How do we resolve?
○ Main thread joins the worker threads and calls the callbacks?
■ Inefficient! Do you know why?
○ Main thread releases GIL before blocking, workers acquire GIL before
callback
Complexity level 2 : Multiple worker threads
Complexity level 2 : Multiple worker threads
● Seems scalable
● C++ dispatching thread doesn’t even need GIL
○ Python is not inherently blocked!
Complexity level 3 : Python background tasks
Complexity level 3 : Python background tasks
def print_packet(packet):
print(packet)
my_network_library = MyNetworkLibrary()
my_network_library.register_callback_on_packet(print_packet)
my_network_library.listenInThread()
while True:
print(get_resource_usage)
[Link](10)

Library is significantly more flexible now!


Complexity level 4 : back and forth
● Your network library maintains some connections to
external processes (distributed computation)
● Python thread is used to interact with those external
processes
○ Perhaps via some scripting
○ Perhaps the user wants to interactively interacting with the external
processes
Complexity level 4 : back and forth
Complexity level 4 : back and forth
● Unsafe: python thread mutating C++ thread managed objects
● Use a lock in the C++ and python thread everywhere?
○ Pros and cons?

● Make C++ thread dispatcher wait for mutation instructions


from python
○ How?
Complexity level 4 : back and forth
template <typename Ret>
Ret safe_dispatch(std::function<Ret(void)>&& fn) {
cpp_chan->dispatch([fn, &done, &excep, &result]() mutable {
try {
result = boost::optional<Ret>{fn()};
} // catch exceptions etc
});
while (!done) { // done should be std::atomic<bool>
usleep(1000);
}
return [Link]();
Library is significantly more flexible now!
}
Complexity level 4 : back and forth
bool MyPythonNetworkLibrary::subscribe(int listener) {
// Thread safe method on network library C++
return internals_->safe_dispatch<bool>(
[this]() { return internals_->subscribe(listener); }
);
}

We have achieved std::numeric_limits<PowerT>::max()


Example 3
Optimizing python
hot loops using *
Some analysis code
reader = [Link](f)
rows = [Link]()
sum_sig_change_by_trade_size, count, last_px = 0.0, 0, 0.0
for i, row in enumerate(rows):
cur_trade_size = 0
for item in [Link]():
id_name = [Link]([Link]) # int => string
... some logic
print("average signal change is " + str(sum_sig_change_by_trade_size / count))
Some analysis code
if id_name == "timestamp":
cur_trade_size = 0
if id_name == "trade_size":
cur_trade_size += [Link]()
if id_name == "signal":
if last_px != 0:
sig_change = abs([Link]() - last_px)
if cur_trade_size > 0:
sum_sig_change_by_trade_size += sig_change / cur_trade_size
count += 1
last_px = [Link]()
cur_trade_size = 0
Hot loop optimization
● User wants to write arbitrary fast analysis
○ Could be written in Cython or Python

● You provide a pre-built .so for parsing files


● Maybe the user wants to write the hot loop themselves?
○ Perhaps we just failed as an API designer 😔?
○ Maybe the problem is just too generic

● Do we just give up?


Let’s look at this again
reader = [Link](f)
rows = [Link]()
sum_sig_change_by_trade_size, count, last_px = 0.0, 0, 0.0
for i, row in enumerate(rows):
cur_trade_size = 0
for item in [Link]():
id_name = [Link]([Link]) # int => string
... some logic
print("average signal change is " + str(sum_sig_change_by_trade_size / count))
“Typed” vs “untyped” For loops
for i, row in enumerate(rows):
cur_trade_size = 0
for item in [Link]():
id_name = [Link]([Link]) # int => string

● Python interpreter interprets some bytecode each time it


hits the given loop
● Perhaps reduce the overhead by “compiling” this code?
Brief digression about cython
● Cython compiles down your py script into C/C++ while
retaining the exact same semantics
● The generated program can also do more things:
○ Deal with actual pointers and C++ data types

● The compiled program keeps most of the performance and


dynamism of an interpreted language, and:
○ is now a C++ .so
○ is not an interpreted script
Cython
/* for item in [Link](): */
__pyx_t_9 = 0;
if (unlikely(__pyx_v_row == Py_None)) {
PyErr_Format(PyExc_AttributeError, "'NoneType' object has no attribute '%.30s'", "items");
__PYX_ERR(0, 17, __pyx_L1_error)
}
__pyx_t_12 = __Pyx_dict_iterator(__pyx_v_row, 0, __pyx_n_s_items, (&__pyx_t_10), (&__pyx_t_11)); if (unlikely(!__pyx_t_12))
__PYX_ERR(0, 17, __pyx_L1_error)
__Pyx_GOTREF(__pyx_t_12);
__Pyx_XDECREF(__pyx_t_6);
__pyx_t_6 = __pyx_t_12;
__pyx_t_12 = 0;
while (1) {
__pyx_t_13 = __Pyx_dict_iter_next(__pyx_t_6, __pyx_t_10, &__pyx_t_9, NULL, NULL, &__pyx_t_12, __pyx_t_11);
if (unlikely(__pyx_t_13 == 0)) break;
if (unlikely(__pyx_t_13 == -1)) __PYX_ERR(0, 17, __pyx_L1_error)
........
Cython
while (1) {
__pyx_t_13 = __Pyx_dict_iter_next(
__pyx_t_6, __pyx_t_10, &__pyx_t_9, NULL, NULL, &__pyx_t_12, __pyx_t_11
);
if (unlikely(__pyx_t_13 == 0)) break;
if (unlikely(__pyx_t_13 == -1)) __PYX_ERR(0, 17, __pyx_L1_error)

● Performance does improve!


● Still has to do a lot of boilerplate to work around the
lack of “type” information
● In my test, it went from 36.3s to 34.2s just in the hot
loop (no IO included)
Cheaper function call conventions - Type informed
for i, row in enumerate(rows):
cur_trade_size = 0
for item in [Link]():
id_name = [Link]([Link]) # int => string

● Writing python, interpreter interprets some bytecode each


time it hits the given statement
○ It needs to handle a lot of possible eventualities
○ We profiled earlier, python function calls are roughly ~100ns
○ Can we inform the runtime that getIdName is always a simple function
call?
● Can Cython do this?
○ No, just strips off the bytecode interpretation code
Pointers..ahem..Capsules

● What if we had a function pointer that we could fetch


early on, and call it in our hot loop?
● What if we could get typed objects out of our python
module, and let cython generate type-informed code?
Pointers..ahem..Capsules

● PyCapsule_New creates a new python object that wraps a


void* pointer and a string description of it.
● PyCapsule_GetPointer takes a capsule object as input and
returns the pointer
○ Only if you provide the same string description

PyObject *PyCapsule_New(void *pointer, const char *name);

void *PyCapsule_GetPointer(PyObject *capsule, const char *name);


Ugly API to expose pointers to python
typedef std::string(get_id_name_ptr)(const void*, int);

boost::python::object getIdNameFunctionPtr() const {


api::v1::get_id_name_ptr* fn(+[](const void* self, int id) -> std::string {
return static_cast<const RowReader*>(self)->getIdName(id);
});
PyObject* result = PyCapsule_New((void*)fn, "get_id_name_ptr_v1", NULL);
return object(detail::new_reference(result));
}
boost::python::object getSelfPtr() const {
PyObject* result = PyCapsule_New((void*)this, "row_reader_ptr", NULL);
return object(detail::new_reference(result));
}
Faster function calls using type information
// user code
get_id_name_ptr = <string(*)(const void*, int) nogil> PyCapsule_GetPointer(
[Link](), "get_id_name_ptr_v1")
self_ptr = <const void*> PyCapsule_GetPointer(
[Link](), "row_reader_ptr")
for i, row in enumerate(rows):
cur_trade_size = 0
for item in [Link]():
item_id = <int>[Link]
id_name = get_id_name_ptr(self_ptr, item_id)
if id_name == string(<const char*>"timestamp"):
cur_trade_size = 0
Capsules
● As expected, performance is better
● 34.2s => 30.5s on a test file with 1 million rows:
○ Had 13000540 function calls
○ Saving ~300ns per function call, adds up!
● Can take this principle further for your highly latency
sensitive applications
○ not useful everywhere, adds ugly pointers to python code
Example 4
open source
libraries
numpy
Optimized numeric operations for python using typed arrays
Brief overview of NumPy
● Abstraction over contiguous and multi-dimensional arrays
exposed in Python (ndarray)
● Contains lots of optimized functions for common
operations over such ndarrays
● Library’s core written in C, the ndarray type and various
common operations on it are implemented in C

In [7]: x = [Link]([6, 7, 8]); print("{}, {}".format(type(x), [Link]))


<class '[Link]'>, int64

In [8]: [Link](x)
Out[8]: 7.0
Brief overview of NumPy
● Source for subsequent slides:
[Link]
[Link]
● Example source code:
[Link]
ultiarray/arrayobject.c
Let’s design a n-dimensional array (brief)
How would you design a library exposing an API to an
n-dimensional fixed size array?

● void* or char* to the start of an allocated memory region


● Store number and length of each dimension
● Store data type size
○ Use that to calculate and store some “strides” for each dimension

● Expose some n-d iterator API for it


Numpy types
typedef struct PyArrayObject {
PyObject_HEAD
char *data;
int nd;
npy_intp *dimensions;
npy_intp *strides;
PyObject *base;
PyArray_Descr *descr;
int flags;
PyObject *weakreflist;
/* version dependent private members */
} PyArrayObject;
Numpy types
● The PyArrayObject is a valid PyObject* so can be
constructed in python using pybinded methods
● Methods on this defined in PyArray_Type struct
● Various other C methods exposed by the library that can
operate on these objects efficiently
● Usually constructed for simple C types (float, double, int,
int64, etc)
Takeaways
● NumPy implemented a complicated datatype and various
methods in C
● Hot path is hidden behind a DSL using (potentially
vectorizable) C methods
○ [Link]
○ [Link]
○ [Link]
● What if we have a custom hot path and we want to loop
over it in user written C++?
○ boost::python::numpy has a nice wrapper around numpy ndarrays!
How do I write a program that operates on numpy arrays?
● Your library needs to interop with some ML library:
○ return a std::vector<float> and convert it to numpy array in python?
● Why not populate numpy arrays in your library?
○ Boost::python has easy constructors

#include <boost/python/[Link]>
namespace bnp = boost::python::numpy;

Py_intptr_t shape[1] = {static_cast<long>(100)};


bnp::ndarray result = bnp::zeros(1, shape, bnp::dtype::get_builtin<T>());
// Example write operation (v is a source range)
std::copy([Link](), [Link](), reinterpret_cast<T*>(result.get_data()));
Pandas
Columnar data manipulation API built on top of numpy
Brief overview of pandas
● Pythonic APIs built on top of numpy to add first class
support for real world data types:
○ Enumerated values (called categoricals)
○ Timestamp values (datetime and timedeltas)
○ Arbitrary python objects

● Supports transformation of n-dimensional arrays:


○ Slicing out certain rows - creating views or a copied buffer
○ Slicing out certain columns - creating views or a copied buffer

● Think of it as a higher level DSL to optimize even more


hot loops than numpy did.
The “how” of pandas
cdef extern from "numpy/ndarrayobject.h":
bint PyArray_CheckScalar(obj) nogil

@[Link](False)
@[Link](False)
def maybe_booleans_to_slice(ndarray[uint8_t, ndim=1] mask):
cdef:
Py_ssize_t i, n = len(mask)
Py_ssize_t start = 0, end = 0
bint started = False, finished = False

for i in range(n):
The “how” of pandas
● Chose to write all code in Cython instead of C / C++

● Benefits:
○ Easy interop, Cython writes the bindings for your helper functions
○ Easy syntax, code looks very similar to python
○ Can import C / C++ constructs easily

● Why aren’t we doing this?


○ Cannot do this with existing C++ code with complex build steps
○ But a pretty good idea for writing new code!

● Always good to know what is out there, even if not C++


Building extension
modules
Going back to numpy
● Any C++ code using numpy has to be built with a version
compatible with the version in the python venv
● Python itself needs to be ABI compatible!
● Else, chaos!

#include <boost/python/[Link]>
namespace bnp = boost::python::numpy;

Py_intptr_t shape[1] = {static_cast<long>(100)};


bnp::ndarray result = bnp::zeros(1, shape, bnp::dtype::get_builtin<T>());
// Example write operation (v is a source container)
std::copy([Link](), [Link](), reinterpret_cast<T*>(result.get_data()));
So how do we release such extension modules to users?
(1) Wheels
● Extension module is built with a minimum-stable-ABI
requirement for CPython headers
● Can stay compatible for future python releases
● Discussion: [Link]
(2) Source builds
● The user builds the extension module with some context of
the python environment it is supposed to run in
● Hard to keep python and C++ build dependencies consistent
(2) Source builds
The versions need to
match to some extent. C++
module needs to be ABI
compatible.
generate environment spec from python venv
● Your C++ library has certain dependencies
○ numpy, python, ...
● The user’s python venv may have some of these libraries
● Bundle a script with your library that parses user’s
python venv
● Creates a build environment which has your C++
dependencies but pinned to versions already present in
the user’s python venv
generate environment spec from python venv
required_build_spec = [
"json={json}",
"llvm={llvm}",
"curl={libcurl}",
"lz4={lz4}",
"mongo-c-driver={mongo-c-driver}",
"numpy={numpy}", ....]

def get_build_spec():
build_spec = []
versions = {entry["name"]: entry["version"] for entry in get_current_environment_spec()}
for b in required_build_spec:
b = [Link](**versions) # some extra stuff needs to be done to handle default versions
build_spec.append(b)
(2) Source builds
Ideally don’t want users to
have to add these into
their own python venvs
static linking C++ dependencies
● Build some libraries statically into your python binding
● Overall just reduces the number of jigsaw pieces we have
to fit together to load the library into the python env
● For cmake:
○ Boost_USE_STATIC_LIBS=ON, Protobuf_USE_STATIC_LIBS=ON
● Some libraries are not that helpful, try this:

if ("${MYLIB_BUILD_STATIC}" STREQUAL "ON")


ADD_LIBRARY(MyLib::MyLib STATIC IMPORTED)
SET_TARGET_PROPERTIES(MyLib::MyLib PROPERTIES IMPORTED_LOCATION "${VENV_DIR}/lib/libmylib.a")
endif()
static linking C++ dependencies
● What if two extension modules use different versions of
your statically linked libblah.a?
○ Potentially conflicting symbols loaded?

● What if:
○ extension module links statically to libblah.a
○ python virtual environment contains different version of [Link]
with potentially conflicting symbols
static linking C++ dependencies
Two common ways of loading .so files into your program:

● RTLD_GLOBAL
○ The symbols defined by this library will be made available for symbol
resolution of subsequently loaded libraries.
● RTLD_LOCAL
○ This is the converse of RTLD_GLOBAL, and the default if neither flag
is specified. Symbols defined in this library are not made available
to resolve references in subsequently loaded libraries.

(copied from dlopen manpage)


static linking C++ dependencies
● Python import uses RTLD_LOCAL to load your dynamically
built extension module
● RTLD_GLOBAL is pretty bad at diamond dependencies
○ Luckily for us we don’t have to deal with this mostly

● Need static linking if:


○ Users’ venv is not expected to have your library
○ Your library links to some library without an explicit rpath and
users’ venv may have a different version

● Need exact version match if:


○ Library is used on the API surface between C++ and Python
Discussion / questions

You might also like