This is your work, valued
Trained as an astrophysicist and data scientist I work for Anaconda OSS leading a number of pydata projects such as fsspec, kerchunk, projspec, akimbo, intake
dask-tutorial-scipy-2018. Materials for "Parallelizing Scientific Python with Dask"
70rfsspec. Rust python FS
36fastparquet. python implementation of the parquet columnar file format.
21async-zarr. Hack wrapper around zarr to make data access async
17dask-gke. kubernetes setup to bootstrap distributed on google container engine
13daskberg. dask client for iceberg (super-alpha)
11splunk_connector. Splunk to dataframes via REST access
8libhdfs3-downstream. a native c/c++ hdfs client (downstream fork from apache-hawq)
8pydata_global_2024. Material for "akimbo: vectorized processing of nested/ragged dataframe columns"
4gcsfs. Pythonic file-system interface for Google Cloud Storage
4misc. Python tidbits that may be useful
3intake. A plugin system for loading your data and making data catalogs.
3dask. Versatile parallel programming with task scheduling
2intake-release-blog. Materials related the release notice of the Intake project
2lance. Modern columnar data format for ML and LLMs implemented in Rust. Convert from parquet in 2 lines of code for 100x faster random access, vector index, and data versioning. Compatible with Pandas, DuckDB, Polars, Pyarrow, with more integrations coming..
1thriftlike. Playing thrift with rust
1fastfeather. Feather to pandas loader without pyarrow
1libhdfs3-feedstock. A conda-smithy repository for libhdfs3.
1blog_fsspec. Article to appear on anaconda.com
1