salesforce/TransmogrifAI là dự án tích hợp AI trên GitHub với 2.3k sao, viết chủ yếu bằng Scala. TransmogrifAI (pronounced trăns-mŏgˈrə-fī) is an AutoML library for building modular, reusable, strongly typed machine learning workflows on Apache Spark with minimal hand-tuning
Tóm tắt dựng từ metadata GitHub của chính dự án — chưa có bài review TopGit. Trang sẽ tự động cập nhật khi bài review đầy đủ được xuất bản.
VÌ SAO CHƯA CÓ REVIEW
TopGit viết bài đầy đủ cho repo có nhiều sao nhất và được yêu cầu nhiều nhất. Trang này là snapshot trong thời gian chờ — xem README gốc ở tab READ ME.
TransmogrifAI (pronounced trăns-mŏgˈrə-fī) is an AutoML library written in Scala that runs on top of Apache Spark. It was developed with a focus on accelerating machine learning developer productivity through machine learning automation, and an API that enforces compile-time type-safety, modularity, and reuse.
Through automation, it achieves accuracies close to hand-tuned models with almost 100x reduction in time.
Use TransmogrifAI if you need a machine learning library to:
Build production ready machine learning applications in hours, not months
Build machine learning models without getting a Ph.D. in machine learning
To understand the motivation behind TransmogrifAI check out these:
Open Sourcing TransmogrifAI: Automated Machine Learning for Structured Data, a blog post by @snabar
Meet TransmogrifAI, Open Source AutoML That Powers Einstein Predictions, a talk by @tovbinm
Low Touch Machine Learning, a talk by @leahmcguire
Skip to Quick Start and Documentation.
Predicting Titanic Survivors with TransmogrifAI
The Titanic dataset is an often-cited dataset in the machine learning community. The goal is to build a machine learnt model that will predict survivors from the Titanic passenger manifest. Here is how you would build the model using TransmogrifAI:
import com.salesforce.op._
import com.salesforce.op.readers._
import com.salesforce.op.features._
import com.salesforce.op.features.types._
import com.salesforce.op.stages.impl.classification._
import org.apache.spark.SparkConf
import org.apache.spark.sql.SparkSession
implicit val spark = SparkSession.builder.config(new SparkConf()).getOrCreate()
import spark.implicits._
// Read Titanic data as a DataFrame
val passengersData = DataReaders.Simple.csvCase[Passenger](path = pathToData).readDataset().toDF()
// Extract response and predictor Features
val (survived, predictors) = FeatureBuilder.fromDataFrame[RealNN](passengersData, response = "survived")
// Automated feature engineering
val featureVector = predictors.transmogrify()
// Automated feature validation and selection
val checkedFeatures = survived.sanityCheck(featureVector, removeBadFeatures = true)
// Automated model selection
val pred = BinaryClassificationModelSelector().setInput(survived, checkedFeatures).getOutput()
// Setting up a TransmogrifAI workflow and training the model
val model = new OpWorkflow().setInputDataset(passengersData).setResultFeatures(pred).train()
println("Model summary:\n" + model.summaryPretty())
Model summary:
Evaluated Logistic Regression, Random Forest models with 3 folds and AuPR metric.
Evaluated 3 Logistic Regression models with AuPR between [0.6751930383321765, 0.7768725281794376]
Evaluated 16 Random Forest models with AuPR between [0.7781671467343991, 0.8104798040316159]
Selected model Random Forest classifier with parameters:
|-----------------------|--------------|
| Model Param | Value |
|-----------------------|--------------|
| modelType | RandomForest |
| featureSubsetStrategy | auto |
| impurity | gini |
| maxBins | 32 |
| maxDepth | 12 |
| minInfoGain | 0.001 |
| minInstancesPerNode | 10 |
| numTrees | 50 |
| subsamplingRate | 1.0 |
|-----------------------|--------------|
Model evaluation metrics:
|-------------|--------------------|---------------------|
| Metric Name | Hold Out Set Value | Training Set Value |
|-------------|--------------------|---------------------|
| Precision | 0.85 | 0.773851590106007 |
| Recall | 0.6538461538461539 | 0.6930379746835443 |
| F1 | 0.7391304347826088 | 0.7312186978297163 |
| AuROC | 0.8821603927986905 | 0.8766642291593114 |
| AuPR | 0.8225075757571668 | 0.850331080886535 |
| Error | 0.1643835616438356 | 0.19682151589242053 |
| TP | 17.0 | 219.0 |
| TN | 44.0 | 438.0 |
| FP | 3.0 | 64.0 |
| FN | 9.0 | 97.0 |
|-------------|--------------------|---------------------|
Top model insights computed using correlation:
|-----------------------|----------------------|
| Top Positive Insights | Correlation |
|-----------------------|----------------------|
| sex = "female" | 0.5177801026737666 |
| cabin = "OTHER" | 0.3331391338844782 |
| pClass = 1 | 0.3059642953159715 |
|-----------------------|----------------------|
| Top Negative Insights | Correlation |
|-----------------------|----------------------|
| sex = "male" | -0.5100301587292186 |
| pClass = 3 | -0.5075774968534326 |
| cabin = null | -0.31463114463832633 |
|-----------------------|----------------------|
Top model insights computed using CramersV:
|-----------------------|----------------------|
| Top Insights | CramersV |
|-----------------------|----------------------|
| sex | 0.525557139885501 |
| embarked | 0.31582347194683386 |
| age | 0.21582347194683386 |
|-----------------------|----------------------|
While this may seem a bit too magical, for those who want more control, TransmogrifAI also provides the flexibility to completely specify all the features being extracted and all the algorithms being applied in your ML pipeline. Visit our docs site for full documentation, getting started, examples, faq and other information.
Adding TransmogrifAI into your project
You can simply add TransmogrifAI as a regular dependency to an existing project.
Start by picking TransmogrifAI version to match your project dependencies from the version matrix below (if not sure - take the stable version):
salesforce/TransmogrifAI có 2.3k sao GitHub — tải lại trang để xem số mới nhất, hoặc xem trực tiếp github.com/salesforce/TransmogrifAI. TopGit phản chiếu số sao của GitHub nhưng không cam kết đến từng phút.
salesforce/TransmogrifAI còn đang phát triển không?
Commit gần nhất trên salesforce/TransmogrifAI là 2 tháng trước (theo timestamp GitHub). Repo có 401 fork — một chỉ báo về mức độ quan tâm của cộng đồng.
salesforce/TransmogrifAI là gì?
salesforce/TransmogrifAI (salesforce/TransmogrifAI) là dự án Scala trên GitHub. Theo mô tả gốc: TransmogrifAI (pronounced trăns-mŏgˈrə-fī) is an AutoML library for building modular, reusable, strongly typed machine learning workflows on Apache Spark with minimal hand-tuning
salesforce/TransmogrifAI so với các dự án AI Tools khác thế nào?
salesforce/TransmogrifAI được TopGit xếp vào nhóm AI Tools, với 2.3k sao GitHub và viết bằng Scala. Xem trang chủ đề AI Tools trên TopGit để so sánh với các dự án tương tự theo số sao và mức độ hoạt động.
salesforce/TransmogrifAI viết bằng ngôn ngữ gì?
salesforce/TransmogrifAI chủ yếu viết bằng Scala. Trường "language" của GitHub dựa trên phần lớn byte ở nhánh mặc định.
Vì sao salesforce/TransmogrifAI được xếp vào nhóm AI Tools?
TopGit xếp salesforce/TransmogrifAI vào nhóm AI Tools dựa trên GitHub topics và mô tả của repo (gắn thẻ: "ai", "automated-machine-learning", "automl"). Việc phân loại dựa trên metadata thật của repo, không phải đoán theo cảm tính biên tập.
Đọc đầy đủ README ở tab phía trên.
Vẫn đang phân vân về TransmogrifAI?
Một cú bấm sẽ gửi câu hỏi kèm trang này cho AI — xem AI nói gì về TransmogrifAI.