apache spark - How to read from a csv file to create a scala Map object?

Question

Welcome To Ask or Share your Answers For Others

apache spark - How to read from a csv file to create a scala Map object?

posted Jan 31, 2022 in Technique[技术] by 深蓝 (71.8m points)

apache spark - How to read from a csv file to create a scala Map object?

I have a path to a csv I'd like to read from. This csv includes three columns: "topic, key, value" I am using spark to read this file as a csv file. The file looks like the following(lookupFile.csv):

Topic,Key,Value
fruit,aaa,apple
fruit,bbb,orange
animal,ccc,cat
animal,ddd,dog

//I'm reading the file as follows
val lookup = SparkSession.read.option("delimeter", ",").option("header", "true").csv(lookupFile)

I'd like to take what I just read and return a map that has the following properties:

The map uses the topic as a key
The value of this map is a map of the "Key" and "Value" columns

My hope is that I would get a map that looks like the following:

val result = Map("fruit" -> Map("aaa" -> "apple", "bbb" -> "orange"),
                 "animal" -> Map("ccc" -> "cat", "ddd" -> "dog"))

Any ideas on how I can do this?

See Question&Answers more detail:os

与恶龙缠斗过久,自身亦成为恶龙；凝视深渊过久,深渊将回以凝视…

1 Reply

深蓝 · Answer 1 · 2022-01-31T07:21:29+0000

scala> val in = spark.read.option("header", true).option("inferSchema", true).csv("""Topic,Key,Value
     | fruit,aaa,apple
     | fruit,bbb,orange
     | animal,ccc,cat
     | animal,ddd,dog""".split("
").toSeq.toDS)
in: org.apache.spark.sql.DataFrame = [Topic: string, Key: string ... 1 more field]

scala> val res = in.groupBy('Topic).agg(map_from_entries(collect_list(struct('Key, 'Value))).as("subMap"))
res: org.apache.spark.sql.DataFrame = [Topic: string, subMap: map<string,string>]

scala> val scalaMap = res.collect.map{
     | case org.apache.spark.sql.Row(k : String, v : Map[String, String]) => (k, v) 
     | }.toMap
<console>:26: warning: non-variable type argument String in type pattern scala.collection.immutable.Map[String,String] (the underlying of Map[String,String]) is unchecked since it is eliminated by erasure
       case org.apache.spark.sql.Row(k : String, v : Map[String, String]) => (k, v)
                                                     ^
scalaMap: scala.collection.immutable.Map[String,Map[String,String]] = Map(animal -> Map(ccc -> cat, ddd -> dog), fruit -> Map(aaa -> apple, bbb -> orange))

Categories

apache spark - How to read from a csv file to create a scala Map object?

apache spark - How to read from a csv file to create a scala Map object?

Please log in or register to add a comment.

Please log in or register to reply this article.

1 Reply

Please log in or register to add a comment.

Just Browsing Browsing

Most popular tags