
۴۴٬۰۰۰تومان
نوع فایل دانلود: EPUB
پس از خرید، یک فایل EPUB دریافت میکنید.
این فایل با Calibre، Apple Books و سایر کتابخوانهای دیجیتال مناسب است.
سیستمهای امروزی پردازندههای چندهستهای و GPU دارند که ظرفیت محاسبات موازی را فراهم میکنند. اما بسیاری از ابزارهای علمی پایتون برای بهرهبردن از این موازیسازی طراحی نشدهاند. با این منبع کوتاه اما کامل، دانشمندان داده و برنامهنویسان پایتون یاد میگیرند که کتابخانه متنباز Dask برای محاسبات موازی چگونه APIهایی ارائه میدهد که موازیسازی کتابخانههای PyData از جمله NumPy، Pandas و Scikit-learn را آسان میکند.
این کتاب را برای دانشمندان داده و مهندسان دادهای نوشتهایم که با پایتون و pandas آشنا هستند و میخواهند مسائلی بزرگتر از توان ابزارهای فعلیشان را مدیریت کنند. کاربران فعلی PySpark میبینند که بخشی از این مطالب با دانش قبلیشان از PySpark همپوشانی دارد؛ با این حال امیدواریم همچنان برایشان مفید باشد — و نه فقط برای فاصلهگرفتن از ماشین مجازی جاوا (JVM).
نویسندگان، هولدن کارائو و میکا کیمینز، نشان میدهند چگونه محاسبات Dask را روی سیستمهای محلی اجرا کنید و سپس برای بارهای کاری سنگینتر به فضای ابری مقیاس دهید. این کتاب کاربردی توضیح میدهد چرا Dask بین متخصصان صنعت و دانشگاهیان محبوب است و سازمانهایی مانند والمارت، کپیتال وان، دانشکده پزشکی هاروارد و ناسا از آن استفاده میکنند.
تمرکز اصلی این کتاب روی علم داده و کارهای مرتبط است؛ چون به نظر ما Dask در همین حوزه بیشترین درخشش را دارد. اگر مسئلهای کلیتر دارید که Dask برایش چندان مناسب به نظر نمیرسد، با کمی سوگیری دوباره پیشنهاد میکنیم کتاب Scaling Python with Ray (انتشارات O’Reilly) را ببینید که تمرکز کمتری روی علم داده دارد.
Dask چارچوبی برای محاسبات موازی با پایتون است که از چند هسته روی یک ماشین تا مراکز داده با هزاران ماشین مقیاس میپذیرد. هم APIهای سطح پایین برای کارها دارد و هم APIهای سطح بالاتر با تمرکز بر داده. APIهای سطح پایینِ کارها، زیربنای یکپارچگی Dask با طیف گستردهای از کتابخانههای پایتون هستند. وجود APIهای عمومی باعث شده اکوسیستمی از ابزارها حول Dask برای کاربردهای گوناگون شکل بگیرد. Continuum Analytics که امروز با نام Anaconda Inc شناخته میشود، پروژه متنباز و تحت حمایت مالی DARPA به نام Blaze را آغاز کرد؛ پروژهای که به Dask تکامل یافت. Continuum در توسعه بسیاری از کتابخانههای ضروری و حتی برگزاری همایشها در حوزه تحلیل داده با پایتون نقش داشته است. Dask همچنان پروژهای متنباز است و بخش زیادی از توسعه آن اکنون با حمایت Coiled انجام میشود. Dask در اکوسیستم محاسبات توزیعشده منحصربهفرد است، چون کتابخانههای محبوب علم داده، محاسبات موازی و محاسبات علمی را یکپارچه میکند. این یکپارچگی به توسعهدهندگان اجازه میدهد بخش زیادی از دانش فعلیشان را در مقیاس بزرگ بهکار بگیرند. آنها اغلب میتوانند بخشی از کدشان را هم با تغییرات اندک دوباره استفاده کنند.
Dask مقیاسپذیری تحلیلها، یادگیری ماشین و سایر کدهای نوشتهشده با پایتون را ساده میکند و به شما امکان میدهد دادهها و مسائل بزرگتر و پیچیدهتری را مدیریت کنید. هدف Dask پر کردن همان جایی است که ابزارهای فعلیتان — مثل DataFrameهای pandas یا پایپلاینهای یادگیری ماشین scikit-learn — بیش از حد کند میشوند (یا اصلاً موفق نمیشوند).
با این کتاب یاد میگیرید:
Dask چیست، کجا میتوانید از آن استفاده کنید و چه تفاوتی با ابزارهای دیگر دارد
چگونه از Dask برای پردازش موازی دستهای داده استفاده کنید
مفاهیم کلیدی سیستمهای توزیعشده برای کار با Dask
روشهای استفاده از Dask با APIهای سطح بالاتر و بلوکهای سازنده
چگونه با کتابخانههای یکپارچهشده مثل scikit-learn، pandas و PyTorch کار کنید
چگونه از Dask با GPUها استفاده کنید
Modern systems contain multi-core CPUs and GPUs that have the potential for parallel computing. But many scientific Python tools were not designed to leverage this parallelism. With this short but thorough resource, data scientists and Python programmers will learn how the Dask open source library for parallel computing provides APIs that make it easy to parallelize PyData libraries including NumPy, Pandas, and Scikit-learn.We wrote this book for data scientists and data engineers familiar with Python and pandas who are looking to handle larger-scale problems than their current tooling allows. Current PySpark users will find that some of this material overlaps with their existing knowledge of PySpark, but we hope they still find it helpful, and not just for getting away from the Java Virtual Machine (JVM).Authors Holden Karau and Mika Kimmins show you how to use Dask computations in local systems and then scale to the cloud for heavier workloads. This practical book explains why Dask is popular among industry experts and academics and is used by organizations that include Walmart, Capital One, Harvard Medical School, and NASA.This book is primarily focused on data science and related tasks because, in our opinion, that is where Dask excels the most. If you have a more general problem that Dask does not seem to be quite the right fit for, we would (with a bit of bias again) encourage you to check out Scaling Python with Ray (O’Reilly), which has less of a Data Science focus.Dask is a framework for parallelized computing with Python that scales from multiple cores on one machine to data centers with thousands of machines. It has both low-level task APIs and higher-level data-focused APIs. The low-level task APIs power Dask’s integration with a wide variety of Python libraries. Having public APIs has allowed an ecosystem of tools to grow around Dask for various use cases. Continuum Analytics, now known as Anaconda Inc, started the open source, DARPA-funded Blaze project, which has evolved into Dask. Continuum has participated in developing many essential libraries and even conferences in the Python data analytics space. Dask remains an open source project, with much of its development now being supported by Coiled. Dask is unique in the distributed computing ecosystem, because it integrates popular data science, parallel, and scientific computing libraries. Dask’s integration of different libraries allows developers to reuse much of their existing knowledge at scale. They can also frequently reuse some of their code with minimal changes.Dask simplifies scaling analytics, ML, and other code written in Python, allowing you to handle larger and more complex data and problems. Dask aims to fill the space where your existing tools, like pandas DataFrames, or your scikit-learn machine learning pipelines start to become too slow (or do not succeed).With this book, you'll learn:What Dask is, where you can use it, and how it compares with other toolsHow to use Dask for batch data parallel processingKey distributed system concepts for working with DaskMethods for using Dask with higher-level APIs and building blocksHow to work with integrated libraries such as scikit-learn, pandas, and PyTorchHow to use Dask with GPUs