Fisseha Berhane continues his RDDs vs DataFrames vs SparkSQL series, this time looking at functions:
Let’s use Spark SQL and DataFrame APIs ro retrieve companies ranked by sales totals from the SalesOrderHeader and SalesLTCustomer tables. We will display the first 10 rows from the solution using each method to just compare our answers to make sure we are doing it right.
All three approaches give the same results, though the SQL approach seems to me to be the easiest.