relational database
This commit is contained in:
@@ -82,6 +82,40 @@ and repository name as key and add that dictionary to db. For example if I am ad
|
||||
packages db, I am actually adding, a Dictionary with ["repotype", foo] to db.
|
||||
This is cheap, bloated and slow. We must have different repository databases for packages.
|
||||
|
||||
* We should really consider using a simple and fast relational database (sqlite) for database operations. Here
|
||||
is the rationale and some issues.
|
||||
- Current implementation uses berkeleydb as a storage of dictionaries, with keys and Pickled Python objects.
|
||||
This approach simplifies coding requirements but leads to some problems.
|
||||
1. Some operations cannot be done effectively. Searching can only be done by iterating all objects and searching
|
||||
one by one. component, reverse dependencies etc are holding lists of package names that are updated on every
|
||||
package operation, Multiple repository problem is
|
||||
|
||||
2.Performance: although fast, pickling and unpickling whole objects is still a time consuming process.
|
||||
Pickled Objects contains object metadata and are larger then the data they hold, so reading, writing pickled
|
||||
Objects to db is slower than reading and writing only necessary fields. large objects (Order of kilobytes)
|
||||
causes berkeleydb data buckets to overflow and caches works less efficently.
|
||||
|
||||
4.Disk usage: Using pickled objects requires more memory and disk space. Total size of a current pisi database
|
||||
is over 100Mb, if we modify files.bdb to only hold filenames and packages, size is still around 50MB.
|
||||
|
||||
Rational Databases offers number of advantages and disadvantages:
|
||||
|
||||
Advantages:
|
||||
+ Relational databases offers much more flexible operations on stored data. Searching, grouping, sorting
|
||||
can be done by database engine efficently.
|
||||
+ There will be only a single database file and much smaller database size compared to berkeleydb files
|
||||
holding pickled python objects.
|
||||
+ Since database no longer holds pickled python objects, it is possible to read and modify the database
|
||||
by other programs written in different languages (Java, C++ etc.)
|
||||
+ SQL lite is generally very fast, probably beating "Berkeleydb with pickled objects" by the order of 4-5 on
|
||||
queries.
|
||||
|
||||
Disadvantages
|
||||
- Using a relational database will require more coding when adding and retrieving contents of objects.
|
||||
- Introducing SQL language makes code more complex and harder to maintain.
|
||||
- Designing a good schema, identifying required indexes etc. takes time
|
||||
- Some developers are superstitious about sql and relational databases ;)
|
||||
|
||||
==> class attributes
|
||||
* In some classes there are some attributes assigned but never
|
||||
used. (see remove_unused_attributes.patch)
|
||||
|
||||
Reference in New Issue
Block a user